05 Aug
|
Cerebras
|
Toronto
Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability. Design cutting-edge solutions for operational challenges.
This position is crucial for enhancing the performance and reliability of AI applications. Your expertise will contribute to automated self-service platforms while mentoring junior SREs to create a robust operational culture. Starting with hands-on experience, you will tackle existing production issues and drive change.
Key Responsibilities:
• Architect scalable solutions for multi-datacenter operations
• Build internal tools for streamlined workflow executions
• Define reliability practices for performance measurement
• Guide mid-level SREs during operational incidents
• Analyze and improve metrics for operational efficiency
Requirements:
• 8+ years in SRE or platform engineering roles
• Strong skills in cluster management and CI/CD systems
• Experience in observability with tools like Prometheus
• Proven capability to lead diverse projects
• Ability to communicate technical strategies clearly
Drive the evolution of reliability practices in AI at Cerebras Systems, shaping the future of technology.
#J-18808-Ljbffr
📌 Senior Staff Site Reliability Engineer (Toronto)
🏢 Cerebras
📍 Toronto