Join Cohere as a Site Reliability Engineer and shape the future of AI systems. Work on high-performance infrastructure in distributed environments in a team-oriented manner.
Cohere’s Model Serving team is on the lookout for an experienced individual who thrives in building scalable AI platforms. This role emphasizes the deployment and optimization of cutting-edge language models, providing efficient API services. Your interaction with customers will be key to customizing deployments that suit diverse needs while ensuring sustainability and high performance.
Key Responsibilities:
• Develop systems to automate service management
• Manage Kubernetes deployment operations effectively
• Enhance observability within various environments
• Participate actively in the on-call duty roster
• Collaborate with internal developers to improve infrastructure
Requirements:
• Minimum of 5 years in production systems engineering
• Proficiency in Kubernetes and multi-cloud solutions
• Solid experience with Linux-based environments
• Strong troubleshooting and teamwork skills
• Knowledge of performance characteristics of accelerators
Join an innovative team committed to building advanced AI technologies with Cohere.
#J-18808-Ljbffr
📌 Experienced Site Reliability Engineer Role (Ontario)
🏢 Visa Hunt
📍 Ontario
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.