Cohere invites applications for a Site Reliability Engineer position, focused on deploying advanced AI solutions. Embrace the challenge of high-availability environments while working remotely.In this role with the Model Serving team, you will play a crucial part in developing automated systems that facilitate the deployment of large language models. Your collaboration with diverse teams will drive improvements in service reliability and performance, ensuring that cutting-edge AI applications are delivered effectively. Create cutting-edge solutions tailored to client needs through effective API endpoint management.Key Responsibilities:
- Automate and manage service operations - Support Kubernetes infrastructure for language models - Drive resilience and observability in systems - Commit to SLOs with regular on-call rotation - Build cross-team relationships to enhance developmentRequirements: - Over 5 years of experience in scalable infrastructure - Expertise in Kubernetes and distributed engineering - Familiarity with major cloud platforms (GCP, AWS, etc.) - Excellent collaboration and problem-solving capabilities - Understanding of resource management and acceleratorsBe a part of Cohere’s mission to transform AI experiences for enterprises.#J-18808-Ljbffr
📌 Cohere Site Reliability Engineer Opportunity (Toronto)
🏢 Visa Hunt
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.