Metrolinx is connecting communities across the Greater Golden Horseshoe. Metrolinx operates GO Transit and UP Express, as well as the PRESTO fare payment system. We are also building new and improved rapid transit, including GO Expansion, Light Rail Transit routes, and major expansions to Toronto’s subway system, to get people where they need to go, better, faster and easier.
Metrolinx is an agency of the Government of Ontario. At Metrolinx, equity, diversity and inclusion are essential to living our values of serving with passion, thinking forward and playing as a team.
I⁢ ONLY: Metrolinx's Innovation and Information Technology group supports female team members via "Go Tech Women" an affinity group for women in Information Technology, led by our Chief Information Officer. If you enjoy technology and innovation, value diversity, appreciate work/balance and are looking for an opportunity to make a better world via public service, Metrolinx would like to hear from you! The SRE and Operations embed resilience, recoverability, and operational excellence across Metrolinx I⁢ services.
Through continuous delivery, infrastructure as code, proactive observability, and tested disaster recovery capabilities, the function reduces operational risk and enables rapid recovery from major outages, cyber incidents, and disruptions.
Drive SRE & Operations Leadership: Lead, coach, and direct engineering teams responsible for day-to-day infrastructure operations, platform reliability, data center management, and disaster recovery. Hands-On Leadership Role: Operate as a hands-on technical leader who remains actively engaged in architecture, engineering decisions, operational problem-solving, incident response, and delivery execution while leading and developing the team. Build and govern the organization’s SRE operating model, priorities, and cloud reliability standards covering availability, scalability, recoverability, and operational support across public and hybrid cloud environments.
Define production readiness criteria and operational acceptance guardrails for new or materially changed services, ensuring architecture, monitoring, logging, alerting, security controls, and recovery automation are fully validated prior to release.
Manage Service Health: Define Service-Level Indicators and Service-Level Objectives for critical services to balance deployment velocity with platform stability.
Direct Incident
Response & Post-Incident Learning: Lead major incident response efforts to restore disrupted services rapidly and facilitate post-incident reviews to ensure root causes are remediated and tracked through engineering backlogs. Day-to-Day Operations:
Oversee day-to-day data centre operational management, including HPE Synergy, Hyper-V, VMware, SAN/NAS storage configuration, F5, NetScaler, file servers, database enablement (Oracle Data Guard, SQL Always On, Oracle ExaCC), data backup, networking integration (Cisco ACI and Firepower, Palo Alto Firewalls) and physical infrastructure tuning.
Proactively Manage Infrastructure Risk: Direct reliability risk assessments, capacity forecasting, and controlled resilience testing including but not limited to multi-region, availability-zone, network, and identity failure scenarios to eliminate single points of failure. Support the governance and allocation of an annual capital budget, managing operational and capital budgets to meet product roadmap objectives within financial constraints. Direct vendor deliverables, evaluate technology proposals, participate in contract negotiations, and establish Service Level Agreements (SLAs) and Operational Level Agreements (OLAs) with strategic partners.
Build high-performance team capabilities in automated operations, infrastructure as code, and SRE disciplines.
Flexible Hours: Preparedness to occasionally work outside regular business hours including evenings and weekends during emergency major incidents or urgent technology deployments.
Education Completion of a degree in Computer Science, Information Systems, Business Administration, Engineering, or a related discipline.
Microsoft Certified Azure Solutions
Expert, Azure DevOps Engineer, or Equivalent is an asset.
Cisco Certified Network
Professional (CCNP) is an asset.
Certified Information System Security
Professional is an asset.
Certified Kubernetes
Administrator is an asset.
Information Technology Infrastructure
Library (ITIL) Foundations is an asset. Agile certification (Agile Certified Skilled, Certified Scrum Product Owner) is an asset.
Project Management
Professional (PMP) or PRINCE2 is an asset.
Professional
Experience & Technical Mastery Demonstrated years of progressively responsible experience in infrastructure operations, platform engineering, cloud engineering, or SRE within a large-scale service delivery environment, including leadership of technical teams and complex operational programs.
Hands-on experience leading large-scale cloud migrations from on-premises environments to Microsoft Azure, including migration planning, architecture, execution, risk management, cutover, and operational transition.
SRE & Cloud Practice: Deep knowledge of Site Reliability Engineering frameworks, dynamic resource management, service quotas, autoscaling, infrastructure as code (IaC), policy as code, and cloud cost optimization.
Relationship Management: Exceptional communication, facilitation, and negotiation skills to brief executive leadership, influence internal stakeholders, manage vendor contracts, and introduce modern methodologies (DevOps, Lean, Agile).
Microsoft
Azure, Copilot, Exchange Online, Teams, SharePoint, etc.
Network & Application Delivery: Palo Alto Firewalls, F5, Citrix NetScaler, Fortinet SD-WAN, and Infoblox.
Microsoft Active
Directory, Zscaler,Microsoft Entra ID, Okta, Microsoft Defender, and CyberArk Privilege Cloud. Grafana, ServiceNow, Terraform, Ansible, and GitHub.
Operating Systems & Management Platforms: Red Hat Enterprise Linux, Windows Server, SCVMM, SCCM, and Red Hat Satellite. Hands-on experience implementing enterprise recovery technologies, including Azure Cloud DR, Cisco ACI multisite, Zerto, Hyper-V, Oracle Data Guard, SQL Always On, F5 application delivery gateway, and related replication solutions. Understanding of project management practices, ITIL frameworks, risk and business impact analysis (BIA), etc.
If you’re excited about working with Metrolinx but your past experience doesn’t quite align with every qualification of this posting, we encourage you to apply. We invite all interested individuals to apply and encourage applications from members of equity-deserving communities, including those who identify as Indigenous, Black, racialized, women, people with disabilities, and people with diverse gender identities, expressions and sexual orientations.
[email protected] be advised that a Criminal Record Check may be required of the successful candidate. For Internal applicants, with the recent implementation of the Internal Mobility Policy, the internal recruitment process has changed for non-union roles.
Candidates must be in their current role for 12 months prior to applying for another role and each applicant must be in good standing (not participating in a Performance Improvement Plan). Please review all provisions of the policy before submitting your application. Should it be determined that any background information provided is misleading, inaccurate or incorrect, Metrolinx reserves the right to discontinue with the consideration of your application.
📌 Senior Manager, I&IT Site Reliability Engineering & Operations (Toronto)
🏢 Metrolinx
📍 Toronto