12 Aug
|
Cerebras
|
Toronto
Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability. Design innovative solutions for operational challenges.
This position is crucial for enhancing the performance and reliability of AI applications. Your expertise will contribute to automated self-service platforms while mentoring junior SREs to create a robust operational culture. Starting with hands-on experience, you will tackle existing production issues and drive change.
Key Responsibilities:
• Architect scalable solutions for multi-datacenter operations
• Build internal tools for streamlined workflow executions
• Define reliability practices for performance measurement
• Guide mid-level SREs during operational incidents
• Analyze and improve metrics for operational efficiency
Requirements:
• 8+ years in SRE or platform engineering roles
• Robust skills in cluster management and CI/CD systems
• Experience in observability with tools like Prometheus
• Proven capability to lead diverse projects
• Ability to communicate technical strategies clearly
Drive the evolution of reliability practices in AI at Cerebras Systems, shaping the future of technology.
#J-18808-Ljbffr
📌 Senior Staff Site Reliability Engineer (Toronto)
🏢 Cerebras
📍 Toronto