About the RoleThis position is for a Senior Cloud Site Reliability Engineer. You will be responsible for the daily operations of Solace Cloud, our market-leading SaaS offering, across leading cloud providers and platforms such as Amazon Web Services, Microsoft Azure, Google Cloud Platform, Kubernetes, etc.What You Will Do:- Ensuring that the Solace Cloud Services are healthy and reliable, and that SLAs are being met- Design and implement our infrastructure tooling, observability, and automation- Contribute to making the production operations more efficient, less error-prone, etc.- Expert-level knowledge in handling production Incidents in production-grade multi-cloud environments according to industry-standard Incident management process- Process handling service requests and provisioning by the customers.- Proven ability to manage customer escalations and drive resolution in mission-critical, high-impact production environments- Work directly with customers to identify, troubleshoot, and resolve operational issues.- Expert debugging knowledge in Linux and Kubernetes to detect operational issues.- Be on-call rotation and provide 24x7 off-hours supportIdeally, You Will Be:- Highly technical, excited by technology, and eager to stay up to date in a rapidly evolving environment.- Expert-level knowledge in Cloud Networking Solutions- Knowledgeable in demonstrating the ability to debug at a system level and resolve incidents in complex cloud-based environments- Expert in Site reliability engineering and Incident response- A strong communicator who can articulate complex technical issues clearly and concisely & get on the phone with customers.- Experienced in SaaS operations and customer-facing technical supportRequired Skills:- Proven expertise with public cloud providers (AWS, Azure, GCP)
services & features- Proven expertise with cloud Kubernetes infrastructure platforms such as AWS Elastic Kubernetes Service, Azure Kubernetes Service, Google Kubernetes Service- Hands‑on experience with Monitoring tools like Datadog, Kibana, Prometheus etc.- Hands‑on experience with Infrastructure Automation using Terraform, Cloud Formation- Hands‑on expertise in debugging production alerts- Expert-level understanding of Linux Operating Systems- Programmer in languages such as Groovy, Python, and Go- Certified Kubernetes Administrator- Certified Cloud Administrator (AWS, Azure, or GCP)Why You’ll Love Working at Solace- Work with brilliance – Our team is packed with some of the sharpest minds in the industry.- Balance matters – We believe work should fit into your life, not the other way around.- Hybrid‑first – Flexibility is built into how we work, so everyone feels included and empowered.- Values‑driven – We live and breathe our core values: craftsmanship, trust, courage, freedom, momentum, humility, and human experience.- Growth mindset – Our training programs are designed to help you level up, fast.- Customer Obsessed – We’re proud of our world‑class customer lineup (we’re not shy about it).- Keep it fun – We’re social, we keep things straightforward, and we know how to have a good time.- Creative culture – We’ve got a great sense of humour and we make cool videos on topics like MITT and this (check them out!).Role Status: EXISTINGExpected Salary Range: Expected salary range for this role is from $120,000 to $150,000. The final offer within this range will reflect the successful candidate’s skills and experience.At Solace, we are committed to a fair, inclusive, and transparent recruitment process.Need accommodations during the hiring process? Just let us know — we’re here to support you.#J-18808-Ljbffr
📌 Senior Cloud Site Reliability Engineer - $120,000 - $150,000 A Year (Ottawa)
🏢 Solace
📍 Ottawa