Site Reliability Engineer, Aiops - C$94,000 - C$110,000 A Year (Toronto)

Site Reliability Engineer, Aiops - C$94,000 - C$110,000 A Year (Toronto)

14 Aug
|
Cognizant
|
Toronto

14 Aug

Cognizant

Toronto

About the roleAs a Site Reliability Engineer, AIOps, you will make an impact by transforming how our production environments detect, respond to, and learn from operational issues. You will own the design and implementation of AI-driven observability pipelines, self-healing automation, and intelligent incident workflows that measurably improve system reliability across the organization. This is a hands‑on engineering role for someone who thrives at the intersection of site reliability engineering, automation, and AI‑powered operations. You will be a valued member of the Cloud Infrastructure and Security team and work collaboratively with the Manager.In this role, you will:Implement and optimize monitoring solutions using Dynatrace, Splunk, and Moogsoft, leveraging AI/ML capabilities such as Davis AI, Splunk ITSI, and Moogsoft AIOps to detect anomalies, predict incidents, and reduce alert noise across distributed systemsDesign and build AI‑powered operational workflows that automate incident detection, root cause analysis, remediation actions, and post‑incident insightsConfigure and manage PagerDuty for intelligent alerting, escalation policies, and automated incident responseBuild self‑healing automation and remediation playbooks using Ansible, Python, and GitHub Actions, triggered by AI‑driven observability eventsApply SRE principles including SLOs, SLIs, and error budgets to improve system reliability and eliminate operational toilBuild and maintain CI/CD pipelines using Git and GitHub Actions that incorporate observability signals, AI‑driven quality gates, and automated rollback workflowsDevelop Python‑based tooling and integrations that connect monitoring platforms, ticketing systems, and automation enginesDocument runbooks, processes, and workflows for knowledge sharing and operational continuityRequired skills8+ years of hands‑on experience with Dynatrace (including Davis AI), Splunk, Moogsoft AIOps, PagerDuty, Ansible, Git and GitHub Actions,



and Python scriptingProven experience leveraging AI/ML features within observability and incident management platforms for event correlation, predictive alerting, and automated remediationStrong understanding of distributed systems, cloud infrastructure, and reliability engineeringExperience with SLO/SLI design, error budgets, and performance optimizationStrong communication skills and ability to collaborate effectively across engineering teamsPreferred skillsExperience with Red Hat OpenShift, Kubernetes, or DockerExposure to LLM‑based automation or generative AI for operational workflowsBackground in ChatOps frameworks or event‑driven architectureExperience mentoring junior engineers or leading technical workstreamsBackground in IT operations or managed services environmentsTotal compensationWe regularly assess market data to ensure we offer a competitive compensation package for our associates. The base salary for this position ranges between CAD $94,000 to $110,000 per year. Where the successful candidate may fall within the range depends on relevant education, work and/or management experience and other business‑related and job‑necessary qualifications. This position is also eligible for Cognizant’s discretionary annual performance‑based bonus, as well as advantages that support your physical, mental and financial wellbeing.Working arrangementsWe believe hybrid work is the way forward as we strive to provide flexibility wherever possible. Based on this role’s business requirements, this is a hybrid position requiring 4 days a week in a client or Cognizant office in Toronto, Ontario. Regardless of your working arrangement, we are here to support a healthy work‑life balance through our various wellbeing programs.Legal and application informationCognizant will only consider applicants for this position who are legally authorized to work in Canada without requiring employer sponsorship, now or at any time in the future.Closing dateApplications will be accepted until April 24, 2026.#J-18808-Ljbffr

📌 Site Reliability Engineer, Aiops - C$94,000 - C$110,000 A Year (Toronto)
🏢 Cognizant
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer, aiops - c$94,000 - c$110,000 a year (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer, aiops - c$94,000 - c$110,000 a year (toronto) / toronto