23 Aug
|
Ju0026M Group
|
Toronto
23 Aug
Ju0026M Group
Toronto
We are hiring for a Senior DevOps / Site Reliability Engineer AI Platform. The role focuses on building, operating, monitoring, and scaling an enterprise AI platform across Development, QA, and Production environments.
A strong fit will have advanced Azure and Kubernetes experience along with strong expertise in observability, dashboards, monitoring, alerting, and automated scaling.
What you bring
- Strong professional experience in DevOps, Site Reliability Engineering, cloud infrastructure, or platform engineering.
- Advanced hands-on experience with Microsoft Azure.
- Strong experience deploying and operating Kubernetes environments.
- Strong knowledge of Docker and containerization technologies.
- Experience supporting containerized applications across Development, QA, and Production environments.
- Advanced experience designing and building dashboards using Grafana, Kibana, Azure Monitor, Application Insights, or comparable tools.
- Solid experience with monitoring, observability, logging, alerting, and operational health checks.
- Experience with application load, infrastructure capacity, performance, and automated scaling.
- Experience with CI/CD pipelines and automated application deployment.
- Experience with Infrastructure as Code tools such as Terraform, Bicep, or ARM templates.
- Strong troubleshooting skills across applications, containers, infrastructure, networking, and cloud services.
What you'll do
- Design, deploy, configure, and maintain infrastructure within Microsoft Azure.
- Deploy and operate containerized applications using Kubernetes.
- Monitor container and cluster health, resource consumption, capacity, and performance.
- Configure scaling policies and develop intelligent scaling approaches based on workload and resource utilization.
- Design and build operational dashboards covering platform health, performance, capacity, errors, latency, and container health.
- Implement monitoring and alerting across infrastructure, applications, containers, integrations, and AI platform services.
- Establish actionable alerts, health checks, anomaly detection, and automated remediation where appropriate.
- Support production platform stability, availability, and operational readiness.
- Investigate platform, deployment, infrastructure, monitoring, and performance issues.
- Participate in root-cause analysis and implement preventative improvements.
- Create operational runbooks and troubleshooting guidance.
Nice to have
- Experience supporting AI, machine learning, data, or high-compute platforms.
- Experience monitoring AI models, inference services, token usage, GPU workloads, API consumption, queues, or model performance.
- Experience implementing automated remediation, predictive monitoring, or AI-assisted platform operations.
- Familiarity with AWS services and cloud operations.
- Experience with Elasticsearch, Log Analytics, OpenTelemetry, Prometheus, or similar observability technologies.
- Experience defining service-level indicators, service-level objectives, and reliability standards.
- Experience with security, identity, secrets management, and cloud governance within Azure.
📌 Cloud DevOps Engineer (Toronto)
🏢 Ju0026M Group
📍 Toronto