Inscreva-se para aceder a todos os recursos do nosso serviço
  • Pesquisa de emprego
  • Favorito
  • Criar um CV
    Novo
  • Salários
  • Alertas de emprego

Reinforcement Learning Lead

Datamentors

Este cargo requer uma presença local. Consulte empregos idênticos a seguir.

Datamentors is a robotics start-up building an agnostic orchestration platform and a family of proprietary robots — quadruped, mobile upper-body and humanoid — assembled at our own factory in Caniçal, Madeira. We also develop custom Vision-Language-Action (VLA) models that give our robots natural-language reasoning and autonomy, deployable across sectors from hospitality and healthcare to defence and logistics. With prototypes running and a clear path to 600 units per year, the engineers joining now will shape both the hardware and the intelligence behind Ardia for years to come. ROLE OVERVIEW Own the post‐training stage of the Ardia humanoid platform — taking pretrained, behaviour‐cloned Vision‐Language‐Action policies and systematically lifting their success rate, precision, and robustness through RL fine‐tuning, reward modelling, and a tight sim‐to‐real loop. This hands‐on leadership role sets the technical direction for RL across the platform, builds the training and evaluation infrastructure, and grows a small team around it. You'll be the person who turns a capable‐but‐inconsistent humanoid into a dependable one — the difference between a research demo and a product. WHAT YOU'LL DO Post-VLA RL fine-tuning — design and run the pipeline that takes a pretrained/imitation-learned VLA policy and improves it with reinforcement learning (online and offline RL, RLHF/RL-from-feedback, preference and reward-model approaches) to lift task success rates and reduce failure modes. Reward design & evaluation — define reward signals, success criteria, and automated evaluation harnesses that actually correlate with real-world manipulation and whole-body performance, and that catch regressions before deployment. Sim-to-real — build and tune the simulation-to-hardware transfer loop (domain randomization, residual policies, real-world fine-tuning) so gains in simulation hold up on the physical robot. Data flywheel — establish the loop that turns robot rollouts and teleop data into better policies: autonomous data collection, filtering, labelling, and continual retraining. Integration across the stack — work with the teams owning the orchestration layer, the VLA policy, and whole-body control so RL improvements compose cleanly rather than fighting the rest of the system. Technical leadership — set the RL roadmap, make build‐vs‐adopt calls on frameworks and tooling, mentor engineers, and represent the work to the wider team and external partners. WHAT WE'RE LOOKING FOR Strong, demonstrable RL expertise: PPO/GRPO and related policy-gradient methods, offline RL, RLHF/preference-based RL, and reward modelling — you understand where each breaks and why. Robot learning / embodied AI experience — ideally hands‐on with VLA models (e.g. OpenVLA, π0, GR00T-class, or comparable manipulation/locomotion policies). Practical sim-to-real experience and fluency with at least one major simulator (Isaac Lab / Isaac Gym, MuJoCo, or similar). Solid ML engineering: PyTorch, distributed training, GPU-efficient pipelines, and the discipline to build reproducible experiments and rigorous evaluation. Evidence of shipping — you've taken a policy from "works in a demo" to "works reliably," not just published benchmarks. Nice to have Direct experience with humanoid or whole-body control (locomotion + manipulation), and with the realities of training on real hardware. Familiarity with VLA post-training specifically — fine-tuning, distillation, or RL on top of large pretrained action models. Comfort working close to perception (vision, point clouds) and to low-level control. Open-source contributions to robot learning or RL frameworks. Experience standing up data-collection / teleoperation pipelines. Publications or applied work at the VLA / robot-learning frontier (ICRA / CoRL / RSS-level). WHY DATAMENTORS Own the post-training layer end to end, with the autonomy to build it your way. The architecture, hardware, and base policies are already in place — your RL work is the leap that turns a capable demo into a dependable product. Hands-on technical leadership: set the RL roadmap, make the tooling calls, and grow a small team around the work. We assess candidates on demonstrated ability, not credentials alone — if you've done the work and can show it, we want to talk. From research demo to dependable product — own it. #J-18808-Ljbffr

Vaga publicada dia atrás
Empregos semelhantes que podem ser interessantes para vocêCom base na vaga Reinforcement Learning Lead em Funchal
  • €2,300 por mês

     ...Join one of the leading hairdressing chains in the Nordic countries and experience working in an international environment! Are you a passionate...  ...! This is your chance to work in an international environment, learn and grow alongside professionals from different nationalities,... 

    EUTALENTS

    Funchal
    18 dias atrás
  •  ...motivated Portuguese -speaking Customer Agent to join an industry-leading digital solutions provider in Greece. In this role, you’ll be...  ...problem-solving skills. Comfortable using computers and willing to learn new software tools. Ability to work in a team-oriented... 

    Mercier Consultancy

    Funchal
    5 dias atrás
  •  ...to become a part of a dynamically growing international company; ~ Challenging projects giving the unique opportunity to grow and learn. Following the EU General Data Protection Regulation No. 2016/679, we inform you that by responding to this announcement, you... 

    Wire IT

    Funchal
    Um mês atrás
  • €975 a €1,200 por mês

     ...bilingual). Have an English level of B1 or higher. Are ready to live and work in Portugal. Are motivated, reliable, and eager to learn. Are open to a multicultural work environment. Are an EU citizen or hold a valid work permit for Portugal. What we offer:... 

    Nomadify

    Funchal
    Um mês atrás
  • €975 a €1,200 por mês

     ...higher. Fully prepared to relocate, live, and work in Portugal. A proactive, dependable individual with a strong willingness to learn. Comfortable working within a diverse, multicultural corporate setting. Must hold EU citizenship or a valid, existing work... 

    Nomadify

    Funchal
    Um mês atrás
  • €975 a €1,200 por mês

     ...internal tools and systems. Ready to work and build a professional career in Portugal. Highly motivated, reliable, and eager to learn new skills. Comfortable working in a diverse, multicultural team setting. Must hold EU citizenship or a currently valid work... 

    Nomadify

    Funchal
    Um mês atrás
  •  ...supportive and innovative workplace to be! We strongly believe in flat hierarchies, a divers and equal company culture and investing in learning & development as an essential to ensure we support our team with all it needs. Join us and be part of a true success story!... 

    Ryanair DAC

    Funchal
    2 meses atrás
  •  ...stories, and comments on major platforms. You'll help identify violations and play a direct role in gathering data to teach machine learning and AI how to understand human culture. The Metaverse Explorer (Virtual Reality Integrity): Step onto the digital frontier!... 

    Next Job Abroad

    Funchal
    12 dias atrás
  •  ...Mercier Consultancy is recruiting an Portuguese -speaking Customer Service Representative on behalf of a leading international company in the tobacco industry. If you are passionate about customer service and want to work in an international environment, this is a... 

    Mercier Consultancy

    Funchal
    5 dias atrás
  •  ...escritas. Responsabilidades De acordo com o Supervisor e Lead tech de Service, as principais tarefas são: Assegurar que...  ...We also aim to give everyone equal access to opportunity. To learn more about our company and life at Vestas, we invite you to... 

    Vestas

    Câmara de Lobos
    10 dias atrás

Deseja receber mais vagas?

Assine e receba vagas semelhantes a Reinforcement Learning Lead. Seja o primeiro a se candidatar!