22 Aug
|
Veeda AI
|
Toronto
Member of Technical Staff - Data
About Us
Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the chance to make an outsized impact from day one.
Responsibilities
- Multimodal Ingest: Build ingest for video, lidar, and robot trajectories on Ray Data and Daft, with GPU decode (NVDEC, DALI) and resharding into WebDataset and Lance layouts that stream sequentially rather than seeking per sample.
- Curation & Filtering: Decide what earns a slot using blur, exposure, and camera-trajectory scoring plus embedding deduplication over cuVS indexes, and prove each filter with a downstream ablation, not a dataset-size delta.
- Annotation & Auto-Labeling: Produce the labels the models need, such as VLM captions, camera pose from feed-forward reconstruction (VGGT, MASt3R), and depth and segmentation pseudo-labels, and hold each to a measured error rate against human review.
- Real & Synthetic Interop: Normalize episodic data across formats such as LeRobotDataset v3, Open X-Embodiment, and RLDS, reconciling action spaces, control rates, and frame timing, and account for the simulated share of every training mixture.
- Provenance, Licensing & Governance: Track license terms, restricted-source flags, and C2PA content credentials at source granularity,
and version datasets as immutable manifests so any checkpoint traces back to the exact bytes that trained it.
Requirements
- Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in large-scale data engineering.
- Built and operated distributed data pipelines (e.g., Ray Data, Daft, Spark) over hundreds of terabytes, with rigor in idempotency, backfills, and schema evolution.
- Strong Python skills and comfortable in the video stack (codecs, containers, ffmpeg, GPU decode) and in columnar and object storage formats.
- Able to design and defend a data mixture empirically, running curation ablations that measure downstream model quality.
- Experience working inside real licensing constraints on what may and may not be trained on, with provenance treated as a hard requirement.
Nice to Have
- Experience with robot trajectory formats and tooling such as LeRobot, RLDS, ROS 2 bags, or MCAP.
- Experience with sensor calibration, hardware time synchronization, and non-pinhole camera models such as fisheye or ftheta.
- Experience running GPU-accelerated curation with RAPIDS or NeMo Curator.
- Experience building PII, face, and plate redaction into a video pipeline at scale.
- Managed annotation vendors and built the QA statistics that keep them honest.
- Built lakehouse storage on Iceberg or Delta and reduced object-storage cost without losing read throughput.
- Published on data curation or contributed to open-source data tooling such as DataTrove or video2dataset.
#J-18808-Ljbffr
📌 Member of Technical Staff - Data (Toronto)
🏢 Veeda AI
📍 Toronto