Reinforcement Learning Engineer - Ingénieur(e) en apprentissage par renforcement (Montreal)

Reinforcement Learning Engineer - Ingénieur(e) en apprentissage par renforcement (Montreal)

11 Aug
|
NBCUniversal
|
Montreal

11 Aug

NBCUniversal

Montreal

Job Description We are seeking a Reinforcement Learning Engineer with experience manipulating virtual environments to train autonomous agents.

This role focuses on the design of robust simulation environments, reward structures, and policy architectures that can navigate complex, multi-sensor landscapes.

Key Responsibilities Cross-Functional Coordination: Work with partner ML and Annotation engineers and TPMs to spec out data, simulation, and training requirements.

Workplace Design: Build and maintain high-fidelity 2D/3D simulation environments (using tools like Unity, Unreal, or Isaac Sim) that serve as the training ground for RL agents.

Reward Engineering: Design and tune complex reward functions that align agent behavior with product goals and safety constraints.

Algorithm Implementation: Develop and optimize RL algorithms (, PPO, SAC, or Offline RL) capable of handling high-dimensional 3D observation spaces.

Sim-to-Real Strategy: Analyze the ''reality gap'' and implement domain randomization or adaptation techniques to ensure models perform reliably in real-world scenarios.

Nous sommes la recherche dun(e) ingnieur(e) en apprentissage par renforcement ayant de lexprience dans la cration et lexploitation denvironnements virtuels pour lentranement dagents autonomes.

Ce rle consiste concevoir des environnements de simulation robustes, des structures de rcompense et des architectures de politiques capables dvoluer dans des contextes complexes et multi-capteurs.

Vous jouerez un rle cl dans le rapprochement entre simulation et performance relle en dveloppant des systmes RL volutifs et en garantissant un comportement fiable des agents dans des conditions varies.

Collaboration interfonctionnelle: Travailler avec les ingnieurs ML, les quipes dannotation et les TPM afin de dfinir les besoins en donnes, en simulation et en entranement.

Conception denvironnements: Dvelopper et maintenir des environnements de simulation 2D/3D haute fidlit laide doutils tels que Unity, Unreal ou Isaac Sim.

Ingnierie des rcompenses:



Concevoir et optimiser des fonctions de rcompense afin daligner le comportement des agents avec les objectifs produit et les contraintes de scurit.

Implmentation dalgorithmes: Dvelopper et optimiser des algorithmes dapprentissage par renforcement (ex. : PPO, SAC, RL hors ligne) adapts des espaces dobservation haute dimension.

Stratgie sim-to-real: Rduire lcart entre simulation et ralit laide de techniques comme la randomisation de domaine et ladaptation afin dassurer des performances fiables en conditions relles.

Qualifications Education: Graduate degree (Masters or PhD) in Robotics, Computer Science, AI, or a related field with a focus on Reinforcement Learning, Imitation Learning, or other Online Machine Learning fields.

Professional Experience: Proven experience as an RL Engineer or Research Engineer in a fast-paced environment.

Industry Context: Prior experience in industries with complex multi-disciplinary teams such as robotics, smart grids, precision agriculture, game development, or aerospace.

Technical Proficiency: Core Tools: Fluency with Python, Git, and the Unix shell. RL Frameworks: Deep familiarity with frameworks like Ray Rllib, Stable Baselines3, or CleanRL.

Physics & 3D Engines: Experience with physics engines (MuJoCo, Bullet) or 3D game engines.

Ecosystem: Familiarity with collaborative tools such as Jira/Confluence, Slack, a Git server, and an experiment tracking framework.

Attributes: Strong Mathematical Background: Essential for understanding Markov Decision Processes (MDPs) and gradient-based optimization.

High Attention to Detail: Critical for debugging non-deterministic agent behaviors and ensuring environment parity.

Formation: Matrise ou Doctorat en robotique, informatique, intelligence artificielle ou domaine connexe avec une spcialisation en apprentissage par renforcement,



imitation ou apprentissage en ligne.

Exprience: Exprience dmontre en tant quingnieur(e) en apprentissage par renforcement ou en recherche dans un environnement dynamique.

Contexte industriel: Une exprience dans des secteurs multidisciplinaires tels que la robotique, les rseaux intelligents, lagriculture de prcision, les jeux vido ou larospatiale est fortement valorise.

Comptences techniques Outils principaux: Excellente matrise de Python, Git et des environnements Unix.

Frameworks RL : Exprience avec des frameworks tels que Ray RLlib, Stable Baselines3 ou CleanRL.

Physique et simulation: Exprience avec des moteurs physiques (MuJoCo, Bullet) ou des environnements de simulation 3D. cosystme: Familiarit avec des outils collaboratifs tels que Jira, Confluence, Slack, les workflows Git et les plateformes de suivi dexpriences.

Qualits recherches Solides bases mathmatiques: Bonne comprhension des processus de dcision de Markov (MDP) et de loptimisation base sur le gradient.

Rigueur et prcision: Capacit dboguer des systmes non dterministes et assurer la cohrence et la prcision des environnements de simulation.

Additional Information As part of our selection process, external candidates may be required to attend an in-person interview with an NBCUniversal employee at one of our locations prior to a hiring decision. NBCUniversal''s policy is to provide equal employment opportunities to all applicants and employees without regard to race, color, religion, creed, gender, gender identity or expression, age, national origin or ancestry, citizenship, disability, sexual orientation, marital status, pregnancy, veteran status, membership in the uniformed services, genetic information, or any other basis protected by applicable law.

If you are a qualified individual with a disability or a disabled veteran and require support throughout the application and/or recruitment process as a result of your disability, you have the right to request a reasonable accommodation.

You can submit your request to AccessibilityS.

📌 Reinforcement Learning Engineer - Ingénieur(e) en apprentissage par renforcement (Montreal)
🏢 NBCUniversal
📍 Montreal

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: reinforcement learning engineer - ingénieur(e) en apprentissage par renforcement (montreal) / montreal

Subscribe to this job alert:

Get the latest job offers by email for: reinforcement learning engineer - ingénieur(e) en apprentissage par renforcement (montreal) / montreal