Join Baseten as a Senior Site Reliability Engineer, enhancing the reliability of AI service infrastructure. Drive observability and automate operational processes to ensure seamless AI performance.
As part of the SRE team, you’ll establish and maintain exemplary standards for day two operations for our multi-cloud Kubernetes systems. Your work will include building observability infrastructures and automating incident management processes. Collaborating closely with engineering teams, you'll turn insights from failures into streamlined solutions for ongoing reliability.
Key Responsibilities
- Own the reliability of our multi-cloud Kubernetes system
- Build comprehensive observability and alerting tools
- Document and improve incident response runbooks
- Automate repeat failure resolutions effectively
- Troubleshoot and resolve model lifecycle challenges
Requirements
- Deep knowledge of Kubernetes environments
- Solid skills in observability tools
- Experience with infrastructure automation
- Capable of navigating engineering and ops intersections
- Interest in AI deployments and scaling
Drive reliable AI performance systems at Baseten with your expertise.
📌 Senior Site Reliability Engineer at Baseten (Montreal)
🏢 Baseten
📍 Montreal
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.