Site Reliability Engineer (Toronto)

Site Reliability Engineer (Toronto)

10 Aug
|
Gemini Solutions
|
Toronto

10 Aug

Gemini Solutions

Toronto

We are looking for a Senior Site Reliability Engineer (SRE) with a strong platformownership mindset to drive reliability, scalability, and performance of mission-critical, distributed systems.This role sits at the intersection of software engineering, cloud infrastructure, and production operations, with a focus on building resilient systems, improving observability, automating operations, and driving reliability at scale.You will act as a technical lead for platform reliability, working closely with engineering and business stakeholders to ensure systems are highly available, performant, and continuously improving.Experience:3+ years of experience in SRE, DevOps, or Production EngineeringExperience working in production-critical environments with high availability requirementsExposure to global systems and cross-team collaborationKey ResponsibilitiesPlatform Reliability & OwnershipOwn availability, performance, and scalability of production systemsDefine and implement SLIs, SLOs, and error budgetsDrive continuous improvements in system resilience and efficiencyIncident Management & Root Cause AnalysisLead end-to-end incident response and service restorationPerform deep root cause analysis across infrastructure, application, data, and network layersImplement long-term fixes and reduce recurrence through engineering improvementsObservability & MonitoringDesign and enhance monitoring, logging,



and alerting systemsDevelop actionable dashboards and improve alert qualityEnable proactive detection of system issuesAutomation & DevOps PracticesAutomate operational workflows to reduce manual effortBuild and maintain CI/CD pipelinesImplement Infrastructure as Code (IaC) for scalable infrastructure managementManage and optimize systems on contemporary cloud platformsTroubleshoot distributed systems across compute, storage, and network layersDiagnose latency, routing, and performance issues in globally distributed environmentsData & Workflow ReliabilityTroubleshoot data pipelines, job failures, and data inconsistenciesPerform data validation and analysisEnsure reliability across data dependencies and workflowsNetworking & Traffic ManagementDiagnose issues related to DNS, HTTP/S, proxies, and load balancingWork with CDN and edge delivery platforms (e.G., Akamai or similar) to optimize traffic routing and performanceStakeholder CollaborationAct as a liaison between engineering teams and business stakeholdersCommunicate system status, incidents, and risks with clarity and contextPartner with cross-functional teams to drive reliability improvementsAI-Driven Reliability (Emerging Focus)Apply AI/ML-driven techniques for anomaly detection, alert optimization,



andpredictive issue identificationLeverage intelligent automation to improve incident response and operationalCore ExpectationsDemonstrates strong ownership of production systems and outcomesIndependently drives incident resolution and follow-throughApplies structured, analytical thinking to complex technical problemsCommunicates effectively in high-impact, production-critical scenariosFocuses on long-term reliability and scalability improvementsTechnical Skills:Programming & AutomationStrong experience in Python for automation and toolingProficiency in shell scripting (Bash)Experience with API-driven and event-driven automationHands-on experience with AWS, Azure, or GCPStrong understanding of cloud architecture, networking, and security fundamentalsInfrastructure as Code using Terraform, CloudFormation, or AnsibleDevOps & CI/CDExperience with Jenkins, GitLab CI, or similar toolsStrong understanding of build, release, and deployment pipelinesObservabilityExperience with Datadog, Splunk, Prometheus, or GrafanaStrong logging, monitoring, and alerting practicesFamiliarity with incident management tools (e.G., PagerDuty)Data & DatabasesStrong SQL skills for troubleshooting and validationUnderstanding of data pipelines and system dependenciesSystems & PlatformExperience with Docker and containerized environmentsExposure to Kubernetes and web servers (e.G., Nginx)OrchestrationExperience with Airflow, Autosys, or similar scheduling toolsNetworking & CDNStrong understanding of DNS, HTTP/S, proxies, and load balancingExperience with CDN and edge delivery platforms (e.G., Akamai or similar) #J-18808-Ljbffr

📌 Site Reliability Engineer (Toronto)
🏢 Gemini Solutions
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (toronto) / toronto