06 Sep
|
General Motors
|
Markham
06 Sep
General Motors
Markham
This means the successful candidate is expected to report to Markham office three times per week, at minimum.
General
Motors is transforming the automotive landscape through its next-generation Software-Defined Vehicle platform. Data is central to that transformation, powering safety, personalization, energy optimization, operational decision-making, and connected customer experiences. We are seeking a Staff Engineer to help make GM's data platforms reliable, observable, operable, and scalable.
This is a senior technical leadership role for someone who can move comfortably between system-level design, production operations, incident response, automation, and customer partnership. You will help define and spread the engineering patterns that make services easier to operate. You will work with SRE, data engineering, infrastructure, developer experience, application, and product teams to improve reliability from design through production and continuously improve how the organization operates.
Lead the design and implementation of scalable, fault-tolerant, and observable infrastructure supporting vehicle telemetry, data ingestion, and platform operations. Lead production readiness efforts across multiple teams - engaging directly in code, shaping reliability standards, guiding architectural improvements, and ensuring applications launch with resilient deployments, strong observability, and predictable operations. Design, implement, and improve CI/CD delivery pipelines that make releases repeatable, safe, observable, and fast.
Establish appropriate quality gates, artifact promotion, deployment verification, progressive delivery, and rollback practices. Partner across SRE, product, and application teams to design and implement meaningful SLOs, SLIs, observability, monitoring and alerting, runbooks, and operational best practices. Build and improve reusable AI workflows, skills, and evaluations.
Apply appropriate validation techniques, including regression testing, structured evaluations, and LLM-as-a-judge approaches where useful. Automate operational work, including incident intake, triage, diagnostics, remediation, evidence collection, service requests, and customer-facing status workflows. Participate in a weekly on-call rotation with 12-hour shifts; the rotation cycles every eight weeks.
Participate in post-incident reviews and drive durable, system-level fixes that prevent recurrence rather than relying on short-term patches or repeated manual workarounds. Partner directly with internal customers to understand their needs, explain technical trade-offs, and improve service outcomes with tact, empathy, and clear communication. Influence technical direction across teams, mentor engineers, and raise engineering standards through design reviews, code reviews, documentation, and hands‑on leadership with cross-functional engineering projects.
Balance reliability, performance, security, delivery speed, and cost when making technical decisions- especially under pressure. 8+ years in SRE, DevOps, or systems engineering, including experience managing or mentoring high-impact teams. ~ Track record of designing and building and maintaining high-scale, cloud-native systems in production (preferably Azure, AWS, or GCP). ~ Hands‑on experience architecting observability patterns, including standardized instrumentation, OTEL collector configuration, SLO/SLI definitions, and deploying observability resources like monitors, alerts, and dashboards. ~ Strong understanding of production readiness, service ownership, SLOs, incident management, post‑incident learning, and continuous reliability improvement. ~ Experience participating in an on‑call rotation and leading technical response to production incidents. ~ Experience designing, operating, and improving CI/CD pipelines. ~ Understanding of GitOps, release strategies, quality gates, deployment verification, progressive delivery, and safe rollback is expected. ~ Strong programming ability in Python, Go, Java, or a comparable language, with disciplined code review, version control, testing, and maintainability practices. ~ Familiarity with AI-assisted software development, LLM application practices, agentic workflows, and evaluation techniques.
~ Ability to work effectively with internal customers, including in difficult or high‑pressure situations, with professionalism, tact, and empathy. ~ Ability to influence without relying on formal authority and to increase adoption of shared engineering patterns. ~ Strong written and verbal communication skills for both technical and non‑technical audiences. ~ BS / MS / PhD in computer science, engineering, physics, mathematics, or another relevant, technical field.
Azure Databricks Azure Event Hubs
Kubernetes configuration management with Helm and Kustomize GitHub Actions, Argo CD, and GitOps-based deployment models LLM application development and testing with Promptfoo, agentic workflows and reusable AI skills with CoPilot Experience operating large-scale data ingestion, processing, and delivery systems such as Fivetran, Apache Flink, Kafka, and Pulsar This is more than an engineering role — it’s an opportunity to shape the future of mobility. With meaningful projects, a collaborative culture, and a global mission, your impact will be tangible and far-reaching. GM DOES NOT PROVIDE IMMIGRATION-RELATED SPONSORSHIP FOR THIS ROLE.
DO NOT APPLY FOR THIS ROLE IF YOU WILL NEED GM IMMIGRATION SPONSORSHIP NOW OR IN THE FUTURE. Paid time off including vacation days, holidays, and supplemental advantages for pregnancy, parental and adoption leave Healthcare, dental, and vision benefits Life insurance plans to cover you and your family Company and matching contributions to a Defined Contribution Pension plan to help you save for retirement General Motors is committed to being a workplace that is not only free of unlawful discrimination, but one that genuinely fosters inclusion and belonging. We strongly believe that providing an inclusive workplace creates an environment in which our employees can thrive and develop better products for our customers.
General
Motors offers opportunities to all job seekers including individuals with disabilities. Join us to help lead the change that will make our world better, safer and more equitable for all by becoming a member of GM's Talent Community (beamery.As a part of our Talent Community, you will receive updates about GM, open roles, career insights and more. #
📌 Staff Engineer, Site Reliability Engineering (Markham)
🏢 General Motors
📍 Markham