Become a vital part of our Technical Operations team as a Systems Reliability Engineer focused on a large-scale enterprise platform. Ensure high availability and performance while implementing advanced monitoring and alerting frameworks.
In this pivotal Systems Reliability Engineer role, you'll take ownership of the Embedded Finance platform's reliability and resiliency. Your expertise with tools like Splunk, Dynatrace, and Grafana will allow you to improve platform health through defined SLOs and incident response leadership. Collaborate with various stakeholders to enhance the service continuity and operational efficiency of our financial services.
Key Responsibilities:
• Own platform reliability, availability, and resiliency
• Design and maintain monitoring frameworks with Splunk and Datadog
• Define and track SLOs, SLIs, and error budgets
• Lead incident response and coordinated remediation efforts
• Advance root cause analysis processes and documentation
Requirements:
• Experience with major monitoring tools such as Splunk and Grafana
• Proven skills in managing large-scale enterprise platforms
• Robust incident response and RCA experience
• Knowledge of reliability engineering principles
• Excellent communication for cross-team collaboration
Elevate service reliability and operational standards as part of our dedicated team.
#J-18808-Ljbffr
📌 Reliable Systems Engineer for Finance (Ontario)
🏢 Mphasis
📍 Ontario
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.