09 Aug
|
Visa Hunt
|
Toronto
Cohere is seeking a Site Reliability Engineer to enhance AI platform performance. Contribute to deploying robust machine learning systems in a remote-friendly environment. As part of the Model Serving team at Cohere, you will develop and operate a platform for large language model API endpoints.
This role requires collaborating with various teams to ensure productive deployment and optimized performance in high-availability environments. You’ll engage with customers to tailor specific deployments that meet their needs, marking a significant impact in AI technology. Key Responsibilities:
- Build self-service systems for managing services
- Automate Kubernetes deployments for language models
- Ensure environment observability and resilience
- Participate in on-call rotation for SLO adherence
- Foster relationships with internal teams for feedback integration Requirements:
- 5+ years in production infrastructure engineering
- Experience with Kubernetes and large distributed systems
- Background in GCP, Azure, AWS, or similar
- Excellent troubleshooting and collaboration skills
- Familiarity with GPUs and distributed system technologies Drive infrastructure excellence and build impactful AI systems as a Site Reliability Engineer at Cohere.
📌 Site Reliability Engineer at Cohere (Toronto)
🏢 Visa Hunt
📍 Toronto