23 Aug
|
Veeda AI
|
Toronto
Member of Technical Staff - ML Performance Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the prospect to make an outsized impact from day one.
Distributed Training Throughput: Own step time and model FLOPs utilization for multi-node video world model training, choosing the tensor, context, and expert parallelism mix in PyTorch FSDP2 and Megatron-Core rather than inheriting a default.
Precision & Numerical Stability: Take BF16, FP8, and NVFP4 recipes from running to converging on Blackwell, chasing scaling-factor and accumulation bugs into the video tokenizer and VAE layers where the activation outliers actually live. compile integration so quadratic attention over long video sequences stops setting step time. Build the detection layer for silent data corruption (SDC), stuck CUDA kernels, and "card-freeze" hangs,
plus asynchronous and tiered checkpointing that makes an interruption cost minutes rather than a day. Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.
Fluency in Python and C++/CUDA with the ability to predict where a kernel will stall from its memory access pattern before profiling.
Experience profiling live training runs with Nsight Systems or the PyTorch profiler and translating traces into quantifiable step-time or MFU improvements.
Experience writing Triton, CUTLASS, or CuTe-DSL kernels, or contributing to open-source kernel libraries.
Experience implementing context or sequence parallelism for long-horizon video or high-token-count models.
Experience running or porting large training workloads on AMD GPUs (ROCm) or Google TPUs (JAX/XLA).
Experience building fault-tolerant training with elastic world size, dynamic node re-queueing, or asynchronous distributed checkpointing. Publications or presentations on machine learning systems, compilers, or high-performance kernels. #
📌 Staff Engineer - Performance Engineering (Toronto)
🏢 Veeda AI
📍 Toronto