Senior or Staff Site Reliability Engineer (Storage Layer Services) (Ontario)

Senior or Staff Site Reliability Engineer (Storage Layer Services) (Ontario)

04 Sep
|
MongoDB
|
Ontario

04 Sep

MongoDB

Ontario

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively current team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently
You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas
You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture
Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs
Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing
Identify and configure key metrics to detect incidents and quantify service health, availability, and performance
Participate in a 24/7 on-call rotation to resolve issues involving the storage infrastructure
Become an expert in infrastructure performance, helping us optimize from the application level all the way to the kernel




Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing)Value efficiency in processes and operationsExpertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or AzurePrefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toilProficiency in Python, Go, or a similar languageHave 6+ years of experience working on software development and operating distributed systemsExperience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-marketHave operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offsPossess a customer-focused mindsetLeading major architectural shifts, such as moving from legacy storage stacks to new multi-tenant storage architectures, including planning and executing large-scale data and workload migrations with tight availability and durability requirementsManaging and scaling infrastructure across multi-cloud environments (AWS, GCP, or Azure)Designing secure, multi-tenant runtime environments at scale

#J-18808-Ljbffr

📌 Senior or Staff Site Reliability Engineer (Storage Layer Services) (Ontario)
🏢 MongoDB
📍 Ontario

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior or staff site reliability engineer (storage layer services) (ontario) / ontario

Subscribe to this job alert:

Get the latest job offers by email for: senior or staff site reliability engineer (storage layer services) (ontario) / ontario