06 Aug
|
Visa Hunt
|
Ontario
Cohere is seeking a Site Reliability Engineer to enhance AI platform performance. Contribute to deploying robust machine learning systems in a remote-friendly environment.
As part of the Model Serving team at Cohere, you will develop and operate a platform for large language model API endpoints. This role requires collaborating with various teams to ensure productive deployment and optimized performance in high-availability environments. You’ll engage with customers to tailor specific deployments that meet their needs, marking a significant impact in AI technology.
Key Responsibilities:
• Build self-service systems for managing services
• Automate Kubernetes deployments for language models
• Ensure environment observability and resilience
• Participate in on-call rotation for SLO adherence
• Foster relationships with internal teams for feedback integration
Requirements:
• 5+ years in production infrastructure engineering
• Experience with Kubernetes and large distributed systems
• Background in GCP, Azure, AWS, or similar
• Excellent troubleshooting and collaboration skills
• Familiarity with GPUs and distributed system technologies
Drive infrastructure excellence and build impactful AI systems as a Site Reliability Engineer at Cohere.
#J-18808-Ljbffr
📌 Site Reliability Engineer at Cohere (Ontario)
🏢 Visa Hunt
📍 Ontario