RBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and platforms that support the wealth management business. Working at the intersection of software engineering, cloud-native operations, observability, and automation, the team plays a central role in delivering reliable digital services for both internal users and clients.
RBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and platforms that support the wealth management business. Working at the intersection of software engineering, cloud-native operations, observability, and automation, the team plays a central role in delivering reliable digital services for both internal users and clients.
As a Senior Site Reliability Engineer, you will bring an engineering-first mindset, strong operational judgment, and a passion for automation to improve system reliability at scale. You will work closely with development, infrastructure, platform, and support teams to build modern observability practices, improve incident response, strengthen reliability engineering standards, and drive the evolution toward intelligent, self-healing operations. This role is ideal for a hands‑on engineer who is equally comfortable improving production resilience, building automation, defining service‑level objectives, and shaping the future of AI‑enhanced operations.
You will help design and implement scalable SRE solutions across the technology estate using tools and platforms such as Elasticsearch, Ansible, GitHub Actions, Dynatrace, PagerDuty, Moogsoft, Kubernetes, OpenShift, Kafka, and emerging AIOps capabilities. Build and enhance the SRE product base, including intelligent monitoring, alerting, reliability testing, anomaly detection, and automated remediation. Design and pilot machine learning‑based anomaly detection capabilities to improve signal quality and move from reactive to predictive operations.
Architect and implement self‑healing solutions that automatically remediate recurring operational issues with appropriate controls and governance. Design human‑in‑the‑loop workflows that balance automation speed with accountability, risk management, and operational oversight. Standardize telemetry and instrumentation across platforms to improve visibility, coverage, and correlation of operational signals.
Contribute to the centralization and evolution of observability and monitoring backends to enable deeper analytics and faster incident triage. Partner with cross‑functional teams to improve monitoring, logging, alerting, incident response, and production readiness practices. Automate operational workflows and platform tasks using Ansible, GitHub Actions, and scripting languages such as Bash, Python, and PowerShell.
Work closely with development teams to understand application changes, production risks, and release readiness, ensuring services meet reliability standards before and after deployment. Support production deployments by advocating for reliability, resilience, and performance improvements. Troubleshoot production issues across application, middleware, infrastructure, and platform layers.
Participate in an on‑call rotation and provide senior operational support for business‑critical systems. Continuously identify opportunities to simplify, automate, and modernize operations using engineering and AI‑driven approaches. 5+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or Systems Engineering roles with strong operational depth. ~ Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience. ~ Strong experience with infrastructure automation and configuration management, particularly Ansible. ~ Strong scripting and automation skills in Bash, Python, PowerShell,
or similar languages. ~ Strong understanding of production operations, incident management, root cause analysis, and reliability engineering practices. ~ Knowledge of cloud-native and distributed systems concepts, including resiliency, scalability, fault isolation, and performance tuning. ~ Understanding of AIOps, AI/ML concepts, or intelligent automation as applied to observability and operations. ~ Ability to work across teams, influence engineering practices, and communicate clearly with technical and non‑technical stakeholders.
Experience with CI/CD and developer platform tools such as Jenkins, Artifactory, and Vault. Familiarity with containerization and cloud platform patterns, including Docker and Kubernetes‑based deployments.
Experience building or operating anomaly detection, predictive alerting, or self‑healing automation solutions. Familiarity with AI governance, model validation, and operational controls in regulated environments.
Experience with reliability testing, chaos engineering, or resilience validation practices. We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.
A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable A world‑class training program in financial services Agile Methodology, Group Problem Solving, IT Systems Integration, Organizational Leadership, Product Services, Software Development Life Cycle (SDLC), System Applications, System Integration Testing (SIT), Systems Software Employment Type: Full time Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and prospect for all. #
📌 Senior Site Reliability Engineer (Remote) (Toronto)
🏢 RBC
📍 Toronto