Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability. Design cutting-edge solutions for operational challenges.
This position is crucial for enhancing the performance and reliability of AI applications. Your expertise will contribute to automated self-service platforms while mentoring junior SREs to create a robust operational culture. Starting with hands-on experience, you will tackle existing production issues and drive change.
Key Responsibilities:
• Architect scalable solutions for multi-datacenter operations
• Build internal tools for streamlined workflow executions
• Define reliability practices for performance measurement
• Guide mid-level SREs during operational incidents
• Analyze and improve metrics for operational efficiency
Requirements:
• 8+ years in SRE or platform engineering roles
• Solid skills in cluster management and CI/CD systems
• Experience in observability with tools like Prometheus
• Proven capability to lead diverse projects
• Ability to communicate technical strategies clearly
Drive the evolution of reliability practices in AI at Cerebras Systems, shaping the future of technology.
J-18808-Ljbffr
📌 Senior Staff Site Reliability Engineer Toronto
🏢 Cerebras
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.