31 Aug
|
Cerebras
|
Toronto
Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability. Design cutting-edge solutions for operational challenges.This position is crucial for enhancing the performance and reliability of AI applications. Your expertise will contribute to automated self-service platforms while mentoring junior SREs to create a robust operational culture. Starting with hands-on experience, you will tackle existing production issues and drive change.Key Responsibilities:Architect scalable solutions for multi-datacenter operationsBuild internal tools for streamlined workflow executionsDefine reliability practices for performance measurementGuide mid-level SREs during operational incidentsAnalyze and improve metrics for operational efficiencyRequirements:8+ years in SRE or platform engineering rolesStrong skills in cluster management and CI/CD systemsExperience in observability with tools like PrometheusProven capability to lead diverse projectsAbility to communicate technical strategies clearlyDrive the evolution of reliability practices in AI at Cerebras Systems, shaping the future of technology.
📌 Senior Staff Site Reliability Engineer (Toronto)
🏢 Cerebras
📍 Toronto