Robotics Perception Engineer, Vision Models & Mapping (Toronto)

Robotics Perception Engineer, Vision Models & Mapping (Toronto)

02 Sep
|
ndimensions labs
|
Toronto

02 Sep

ndimensions labs

Toronto

Full time position

- Boston, MA or Toronto, ON

About Ndimensions Labs At Ndimensions, we're inventing the infrastructure for next-generation robotics AI systems. Almost everything our robots know about the world starts as pixels from a camera, and so do most of the hard problems we have left. We're looking for a Robotics Perception Engineer to own that layer. This role sits between our AI and navigation teams, working on the vision problems both depend on: whether the robot is really picking up the object it was asked to, whether it can find an object described in plain language, whether perception is fast enough to run on the robot itself, and whether the cameras are calibrated well enough to trust. It is a hands-on role on real hardware in real homes.

What You'll Do

- Close the loop between language and pixels: verify that the object the robot grasps is the object it was commanded to grasp, both in recorded training data and live at inference time.
- Ground language in the scene, turning a text query and an RGB-D frame into a pick target or a navigation goal.
- Take our maps from geometric to semantic: build a persistent, object-level representation of a home that the robot can query in language, so a named object resolves to a place it can drive to.
- Give the map a sense of what moves. Classify people, pets, and movable furniture as dynamic or semi-static, so the robot knows which parts of a map to trust and which to expect to have changed.
- Evaluate, optimize, and deploy detection and segmentation models on-robot within a real-time latency budget on embedded GPUs.
- Build vision-based quality checks over our training data: pre-verify episodes, catch mis-routed or dropped camera streams, and redact faces and personal information from what we keep.
- Own camera calibration and rectification across the robot, and the metrics that show depth and localization are good enough to trust.

What We're Looking For
- Strong background in computer vision for robotics,



with real experience on physical systems rather than datasets and simulation alone.
- Experience with vision-language models: grounding text to pixels, and judging whether a model's output actually matches the instruction it was given.
- Experience with semantic or open-vocabulary 3D scene representations: fusing 2D detections into a persistent map and keeping object identity stable across viewpoints and sessions.
- Experience training, fine-tuning, and deploying detection and segmentation models, and optimizing them for real-time inference on embedded hardware.
- Deep hands-on calibration experience and command of the underlying geometry: intrinsics, stereo and multi-camera extrinsics, rectification, hand-eye calibration, projection and back-projection, and frame conventions.
- Practical experience with depth from cameras, classical or learned, and a feel for how depth error propagates downstream.
- Solid Python and working C++, with ROS2 experience on real robots.
- A measurement-first instinct: you define the metric and build the rig before claiming an improvement, and you can design a credible evaluation when no ground truth exists.
- Master's or PhD in Computer Vision, Robotics, Computer Science, or a related field (or equivalent industry experience).

Bonus (Not Required)
- Publications in vision or robotics venues (CVPR, ICRA, IROS, CoRL), or open-source contributions to vision and calibration tooling.
- Work on vision-language-action policies or other multimodal models for manipulation.
- Experience with visual or visual-inertial odometry and SLAM, or with scene graphs and other queryable map representations.
- Experience building or evaluating multimodal perception for vision-language-action (VLA) policies, including grounding observations and instructions into objects, locations, or manipulation targets.

Apply for this Position We're looking for engineers who want robots to see reliably in real homes, not just on benchmarks.

Please include links to relevant work (papers, repos, demos) in your application.

← Back to All Positions

📌 Robotics Perception Engineer, Vision Models & Mapping (Toronto)
🏢 ndimensions labs
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: robotics perception engineer, vision models & mapping (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: robotics perception engineer, vision models & mapping (toronto) / toronto