Hiring: Senior Cloud DevOps / SRE Engineer – Azure (Onsite)
Location: Toronto, ON (Onsite)
We’re looking for a Senior DevOps / Site Reliability Engineer to help build, operate, monitor, and scale a new enterprise AI platform hosted primarily on Microsoft Azure.
This is a hands-on opportunity to establish contemporary DevOps/SRE practices from the ground up, with a strong focus on observability, intelligent dashboards, monitoring, automated scaling, and AI-assisted platform management.
Key Skills:
- Advanced Microsoft Azure experience
- Strong Kubernetes & Docker experience
- Grafana, Kibana, Azure Monitor, Application Insights
- Monitoring, observability, logging & alerting
- Automated scaling, performance & capacity management
- CI/CD pipelines
- Infrastructure as Code – Terraform, Bicep, ARM
- Production reliability & troubleshooting
- SLI/SLO and reliability engineering practices
Nice to Have:
- AI/ML or data platform experience
- GPU, inference, token usage or AI workload monitoring
- Prometheus,
Elasticsearch, Log Analytics, Open Telemetry
- AWS experience
- Automated remediation or predictive monitoring
- Azure security, identity, secrets & governance
Why This Role?
The platform is currently focused at approximately 10–30 containers, providing an opportunity to build the foundation correctly without the pressure of a large-scale production environment. As adoption grows in 2027 and beyond, you’ll have the opportunity to expand the platform’s reliability and operational capabilities.
If you’re a hands-on Senior DevOps/SRE professional who enjoys cloud engineering, observability, intelligent automation, and building modern AI platforms, we’d love to hear from you.
? Interested? Apply or DM me with your updated resume.
[email protected] /
[email protected]
Contact No - (phone hidden)
📌 Senior Cloud DevOps / SRE Engineer – Azure (Ontario)
🏢 Aptino
📍 Ontario