01 Aug
|
Envision Technology Solutions
|
Greater Toronto Area
01 Aug
Envision Technology Solutions
Greater Toronto Area
Job Title: AI Site Reliability Engineer
Location: Preference Greater Toronto Area OR Remote (Canada – EST)
Type: Contract (C2C/W2) Experience Level: Senior (10+ years SRE/Cloud/AI platform operations preferred)
:
Role Summary
We are seeking an AI Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of AI and cloud platforms. This role blends SRE principles, observability, incident management, infrastructure automation, and AI platform operations.
Key Responsibilities
- Build, deploy, and operate cloud services to defined SLOs
- Monitor, troubleshoot, and optimize AI platform reliability and performance
- Implement observability, alerting, automation, and operational excellence practices
- Drive infrastructure automation using Terraform and CI/CD (GitHub Actions)
- Operate and support AI/data platforms (Databricks, ML workspaces, Azure AI Foundry, AWS Bedrock, GCP Vertex AI)
- Collaborate with Data Science and Engineering teams to resolve platform and performance issues
- Participate in on‑call rotations for production services (minor share of role)
- Ensure secure and reliable data movement across pipelines
- Triage performance and reliability issues with vendors (e.g., Databricks)
Required Skills
- Solid SRE/infrastructure engineering background
- Strong Azure experience; working knowledge of AWS and GCP (multi‑cloud)
- Hands‑on with Terraform, Kubernetes, GitHub Actions CI/CD
- Awareness of LLM Ops and AI/ML workflows (model runtimes, fine‑tuning, evaluation)
- Experience operating Databricks or ML workspaces
- Understanding of data security, reliability, and secure data movement patterns
- Solid incident response, RCA, and problem‑solving skills
📌 AI Site Reliability Engineer – Remote-Greater Toronto Area- Canada (EST, Contract,C/
🏢 Envision Technology Solutions
📍 Greater Toronto Area