26 Sep
|
Cerebras
|
Toronto
Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability. Design cutting-edge solutions for operational challenges.
This position is crucial for enhancing the performance and reliability of AI applications. Your expertise will contribute to automated self-service platforms while mentoring junior SREs to create a robust operational culture. Starting with hands-on experience, you will tackle existing production issues and drive change.
Key Responsibilities: • Architect scalable solutions for multi-datacenter operations • Build internal tools for streamlined workflow executions • Define reliability practices for performance measurement • Guide mid-level SREs during operational incidents • Analyze and improve metrics for operational efficiency
Requirements: • 8+ years in SRE or platform engineering roles • Strong skills in cluster management and CI/CD systems • Experience in observability with tools like Prometheus • Proven capability to lead diverse projects • Ability to communicate technical strategies clearly
Drive the evolution of reliability practices in AI at Cerebras Systems, shaping the future of technology. #J-18808-Ljbffr
📌 Senior Staff Site Reliability Engineer (Toronto)
🏢 Cerebras
📍 Toronto