Site Reliability Engineer - Python (Vancouver)

Site Reliability Engineer - Python (Vancouver)

11 Sep
|
Socket.dev
|
Vancouver

11 Sep

Socket.dev

Vancouver

Active in over 60 countries and all 50 US States, we protect the customers of the world’s largest digital companies, including Klarna, Revolut, Stripe, Priceline, Agoda, Booking.com, Turkish Airlines, Tongcheng Travel, eBay, and Uber, with seamless, end-to-end experiences.

Cover

Genius has protected more than 73M customers globally across 240M policies with USD $3.2BN in gross written sales. As part of our team, you’ll help drive our AI-first roadmap, developing hyper-personalization engines, agentic distribution, and automated claims infrastructure, while building the scalable technology powering the fast-growing $70B embedded protection market.

Our people are: Accountable, customer-obsessed, collaborative, driven As a Site Reliability Engineer, you'll lead reliability and infrastructure projects that span multiple teams and business domains. To drive success in this role, you will have a strong background in cloud infrastructure and platform engineering, with experience across infrastructure-as-code, CI/CD and release automation, observability, security, and disaster recovery. You should possess strong technical judgement, the ability to independently lead complex projects from design through delivery, and a proactive approach to eliminating operational risk before it becomes a problem.

Regular collaboration with software engineering teams, security teams, and other relevant stakeholders will be key in ensuring the reliability and efficiency of our production systems are achieved. Analyze, test, and evolve systems to improve reliability and performance at an architectural/infrastructure level, leading medium-to-large projects from design through delivery Apply AWS and GCP expertise to architect and build reliable, highly-available cloud infrastructure across multiple products and projects Act as incident commander for significant production incidents, lead troubleshooting on complex issues,



using blameless post-mortems to drive continuous improvement Reduce operational toil by building automation and self-service tooling, that other engineers can adopt, rather than absorbing repetitive work yourself Apply AI-assisted development to infrastructure problems, and guide other engineers on using AI tools effectively within your team's workflows Contribute to capacity planning and cost optimisation Mentor other engineers on systems thinking, incident management, and production ownership, raising the bar within your team Implement security best practices in infrastructure development and maintenance - such as least-privilege access, secrets management, policy-as-code gates in your pipelines 3+ years of experience in SRE, Platform Engineering, DevOps or other related roles ~ Strong understanding of SRE and platform engineering principles, with experience applying them to lead cross-team projects ~ Comfortable scripting and developing internal tooling with Bash and at least one programming language (e.g. Python, Go) ~ Fluent with AI-driven development environments like Cursor, Claude Code, or Codex, with a proven ability to leverage these tools within production engineering workflows ~ Experience working with Linux ~ Strong understanding of networking, distributed systems, and system architecture at scale ~ Proven experience deploying, scaling, and monitoring web applications and databases in high-availability environments ~ Deep expertise in AWS and/or GCP platforms ~ Bachelor's degree in Computer Science/Engineering,



a postgraduate degree and/or record of academic achievement is also desirable Takes ambiguous problems and drives them to shipped outcomes - not just code, but results, and takes accountability even without a clear owner Balances speed with quality — knows when to iterate quick and when to invest in durability Manages risk proactively — identifies failure modes and mitigates before they bite Communicates technical concepts clearly to engineers, product, and business stakeholders AI-First Mindset Treats AI tools as essential infrastructure, not optional add-ons — continuously experiments with new capabilities Helps others adopt AI workflows, shares what works, and raises the floor for the whole team We act decisively, move with intention, and hold a high bar.

Champion Our Customers: We lead with empathy and center the customer, turning complex disruptions into seamless moments of trust.

Flexible Work

Environment - Our teams are hybrid. We work from home on a Wednesday and Thursday and attend the office on Monday, Tuesday and Friday with flexibility around start/finish times. The Legal & Privacy Stuff Cover Genius promotes diversity and inclusivity.

We don't tolerate discrimination, demeaning treatment of anyone, or harassment due to race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or any other legally protected status. By submitting your application, you acknowledge that we may collect, store, and process your personal data for recruitment purposes. To ensure a fair evaluation, we may use AI to assist in sorting applications, but all final decisions are made by our hiring team and no candidate dispositions are automated.

For detailed information about how we handle your data and our use of AI, please review our full Privacy Policy. *

📌 Site Reliability Engineer - Python (Vancouver)
🏢 Socket.dev
📍 Vancouver

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer - python (vancouver) / vancouver

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer - python (vancouver) / vancouver