25 Aug
|
Cerebras Systems
|
Toronto
25 Aug
Cerebras Systems
Toronto
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.
We are looking for an Inference Platform SDET to join the Inference Service Quality team at Cerebras and work on the inference platform. This team sits at the intersection of distributed systems, cloud and cluster infrastructure, and the software stack that serves the world's fastest AI inference. In this role, you will own the quality and reliability of the infrastructure that deploys and runs the Cerebras Inference Platform — from CI/CD pipelines and Kubernetes-based deployments to ingress, load balancing, and service discovery.
You will validate the platform both in cloud environments and on real Cerebras clusters, working side by side with the Inference Platform development team to catch issues before our customers do. This is an excellent opportunity for engineers who enjoy infrastructure, automation, and debugging across the full deployment stack, and who want to ensure that a platform serving inference at massive scale stays fast, reliable, and production-ready. Design, build, and maintain test infrastructure and automation for deploying and validating the Cerebras Inference Platform.
Validate the platform across environments — from cloud-managed Kubernetes to deployments running on Cerebras hardware. Test and verify deployment infrastructure including Kubernetes workloads, CI/CD pipelines, ingress and service discovery, NGINX, and load balancing.
Collaborate closely with the Inference Platform development team to ensure current features and platform capabilities ship reliably.
Investigate and debug complex issues spanning networking, orchestration, deployment, and distributed services. Develop and maintain testbeds used to validate platform performance, scalability, and reliability. Identify failure points, bottlenecks, and edge cases that impact platform stability and inference performance.
Contribute to test plans and validation strategies for new platform features and releases. Partner with engineering teams to ensure high-quality, production-ready releases of the Cerebras Inference Platform. 3+ years of experience in software engineering, QA/quality engineering, systems engineering, or infrastructure development. ~ Strong programming skills in Python and/or Go (experience with both is a plus). ~ Experience building automation tools, testing frameworks, or internal developer tooling. ~ Experience debugging complex systems, distributed services, or networked infrastructure. ~ Hands‑on experience with Kubernetes and container orchestration in a real production or staging environment.
Experience with GitOps/deployment tooling (e.g., Exposure to performance debugging, profiling, or system observability tools.
Experience with ML inference infrastructure, model serving systems, or GPU-accelerated workloads Inference Service Quality People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. Build a breakthrough AI platform beyond the constraints of the GPU. # Publish and open source their cutting-edge AI research. # Work on one of the fastest AI supercomputers in the world. # Our simple, non-corporate work culture that respects individual beliefs.
We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
📌 Senior SDET, Inference Platform (Toronto)
🏢 Cerebras Systems
📍 Toronto