Infrastructure Engineer - Platform (Toronto)

Infrastructure Engineer - Platform (Toronto)

17 Aug
|
S27a
|
Toronto

17 Aug

S27a

Toronto

Peripheral is developing spatial intelligence, starting in live sports and entertainment. Our models generate spatial data, used for advanced sports analytics and immersive media experiences. We’re solving key research challenges in 3D computer vision, creating the foundations for the next generation of robotic perception and embodied intelligence.

Our team includes engineers and researchers from leading technology companies and research institutions, and we’re building technology at the intersection of AI, graphics, and the future of live entertainment. We’re seeking an experienced Infrastructure Engineer to architect and build the cloud infrastructure that powers Peripheral, from large-scale foundation model training to reliable inference during live events. This is a highly architectural role that requires strong systems thinking and deep experience in DevOps and MLOps.

You’ll help define how systems across Peripheral work together to deliver spatial experiences to customers. At the same time, you’ll build the infrastructure that keeps our research velocity high, including setting up and maintaining ML training clusters, distributed training infrastructure with PyTorch, and optimized multi-view video data loading. You’ll build working knowledge across capture systems, robotics, foundation model training, and production inference to support teams across Peripheral while keeping customer-facing systems reliable.

We’re also looking for someone who can mentor others as the team grows, maintain clear documentation in a fast-moving environment, and ship high-quality infrastructure with strong attention to detail. You’ve built and scaled production cloud infrastructure and are comfortable owning systems end to end, from large-scale training workloads to reliable model serving. Cost, latency, reliability, security, and developer productivity are all first-class concerns,



and you know how to balance long-term architecture with pragmatic execution.

You collaborate well across research and engineering, communicate clearly, value robust documentation, and can provide technical leadership and mentorship as the infrastructure function grows. You’re excited to help shape Peripheral’s infrastructure culture from the ground up, establishing the tooling, standards, and practices the broader engineering organization will build on as the company scales. Design, build, and operate cloud infrastructure for large-scale ML training, including GPU/TPU compute, job orchestration, and experiment tracking.

Partner with research to scale foundation model training workflows, including distributed PyTorch training and efficient multi-view video data loading. Build and operate production infrastructure for model deployment and serving, with a focus on reliability, latency, scalability, and cost. Design reliable data pipelines for moving multi-view video, LiDAR, and other sensor data from capture through customer-facing outputs.

Work across research and engineering to translate infrastructure needs into scalable, self-serve systems and developer tooling. Improve the reliability and resilience of customer-facing systems during live, time-sensitive workloads. Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent practical experience. ~4+ years of experience building and operating production cloud infrastructure, ideally for ML workloads. ~ Deep expertise in AWS or GCP, with the ability to work effectively across both.



~ Experience building distributed training infrastructure, including GPU/TPU clusters, job scheduling, orchestration, and large-scale training workloads. ~ Experience deploying and operating production ML inference systems with strong requirements around reliability, latency, scalability, and cost. ~ Strong understanding of cloud security, networking, observability, and operational best practices. ~ Comfortable working cross-functionally across research, product, and hardware teams that generate and consume data. ~ Must have the legal right to work in Canada and be willing to relocate to Toronto for an in-office role.

We are unable to provide immigration sponsorship at this time.

Experience with AWS and GCP services such as S3, EC2, ECR, Batch, ParallelCluster, SageMaker, GCS, Cluster Toolkit, or equivalent infrastructure.

Experience with ML infrastructure and tooling such as Weights & Biases, MLflow, CVAT, and Hugging Face.

Experience with high-throughput ingestion or streaming systems such as Kafka, Pub/Sub, or Kinesis, particularly for video, point clouds, or other large sensor data.

Experience with real-time video, low-latency streaming, CDNs, or large-scale data delivery systems.

Experience optimizing ML systems, including data loading, distributed training, GPU utilization, or custom kernels. Prior ML research or ML systems experience, including publications, patents or open-source contributions.

Experience as an early or founding infrastructure or platform engineer at a startup.

Experience mentoring junior engineers or intern Competitive equity package as an early team member. Annual salary of $150K–$200K CAD plus performance bonuses, commensurate with experience. Full ownership of high-impact projects shaping the future of spatial intelligence and 3D media.

Adaptable Paid Time

Off (PTO). Comprehensive health, dental, vision, and wellness benefits. #

📌 Infrastructure Engineer - Platform (Toronto)
🏢 S27a
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: infrastructure engineer - platform (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: infrastructure engineer - platform (toronto) / toronto