06 Aug
|
Cerebras
|
Toronto
Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability. Design creative solutions for operational challenges.This position is crucial for enhancing the performance and reliability of AI applications. Your expertise will contribute to automated self-service platforms while mentoring junior SREs to create a robust operational culture. Starting with hands-on experience, you will tackle existing production issues and drive change.Key Responsibilities:
- Architect scalable solutions for multi-datacenter operations
- Build internal tools for streamlined workflow executions
- Define reliability practices for performance measurement
- Guide mid-level SREs during operational incidents
- Analyze and improve metrics for operational efficiencyRequirements:
- 8+ years in SRE or platform engineering roles
- Strong skills in cluster management and CI/CD systems
- Experience in observability with tools like Prometheus
- Proven capability to lead diverse projects
- Ability to communicate technical strategies clearlyDrive the evolution of reliability practices in AI at Cerebras Systems, shaping the future of technology.#J-18808-Ljbffr
📌 Senior Staff Site Reliability Engineer (Toronto)
🏢 Cerebras
📍 Toronto