18 Sep
|
Appit
|
Quebec City
APPIT Software Solutions in Montreal is seeking a Reinforcement Learning Engineer to design RL systems for enterprise optimization, building adaptive agents and RLHF alignment of large language models.
You will implement RL algorithms (PPO, SAC, DQN, MCTS), build simulation environments for training and evaluation, and collaborate with research teams to translate advances into production applications in a quick-paced AI product company.
#J-18808-Ljbffr
📌 Senior RL Engineer - Optimization & LLM Alignment (Quebec City)
🏢 Appit
📍 Quebec City