Manager of Site Reliability Engineering (British Columbia)

Manager of Site Reliability Engineering (British Columbia)

04 Sep
|
Mastercard
|
British Columbia

04 Sep

Mastercard

British Columbia

The Business Operations team is seeking a highly motivated and experienced Manager, Site Reliability Engineering (SRE) to join our team
You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power Mastercard’s global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor
The role of the Business Operations Site Reliability Engineer is to be the production readiness steward for Mastercard products. As Business Operations SRE, we are responsible for ensuring that our platform is secure and healthy
We break down barriers to running our products by fostering developer run ownership and empowering developers to build resilient products
We support our developers during the application build phase in software run principles that include operational design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture
Oversee a team of individual contributors, supporting the execution of strategic initiatives by providing technical expertise and leadership within the Site Reliability Engineering discipline to analyze complex problems and provide novel solutions and/or improvements.
Guide the team in automating routine tasks, troubleshooting complex issues, and optimizing system performance.
Collaborate with cross-functional teams to develop strategies for system scalability and resilience, training team members on technical skills, operational best practices, and incident management.
Oversee incident response efforts, ensuring timely resolution and comprehensive root cause analysis
Cultivate a culture of continuous improvement by promoting best practices, innovation,



and proactive risk management.
Support the implementation and maintenance of high-availability systems to ensure operational stability.
Contribute to documentation, knowledge sharing, and best practices to improve team operational procedures.
Lead automation and scripting efforts to streamline operational processes and incident response workflows.
Manage a team of individual contributors(s) and/or technical lead(s), directing area processes and work to ensure that they align with functional best practices and organizational standards; conduct goal setting and performance appraisal processes to coach team members and support their professional development
Serve as the primary contact responsible for the overall application health, performance, and capacity
Support services before they go live through activities such as system design consulting, capacity planning and launch reviews
Partner with the development and product team of a new application to establish the right monitoring and alerting strategy and create the framework to achieve zero downtime during deployment
Serve as the primary contact responsible for ensuring application scalability, performance, and resilience
Practice sustainable incident response and blameless post-mortems while taking a holistic approach to problem-solving and optimizing time to recover
Automate data-driven alerts to proactively elevate issues. Work with development teams to establish SLOs and improve reliability
Tackle complex development, automation,



and business process problems. Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation, and refinement
Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead Mastercard in DevOps automation and best practices
Increase automation and tooling to reduce toil and manual interventiono Analyses ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns
Willingness and ability to learn and take on challenging opportunities and to work as a member of a matrix-based, diverse and geographically distributed project teamAbility to balance doing things right with fixing things quickly. Flexible and pragmatic, while working towards improving the long-term health of the systemBS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics), or equivalent practical experienceComfortable collaborating with cross-functional teams to ensure that expected system behaviour is understood and monitoring exists to detect anomalies.Experience with DevOps practices and tools such as Chef, Ansible, Artifactory, GitHub, Bitbucket, Jenkins, XLR, and RemedySystematic problem-solving approach, coupled with strong communication skills and a sense of ownership and driveInterest in designing, analyzing, and troubleshooting large-scale distributed systemsAppetite for change and pushing the boundaries of what can be done with automation. Be curious about new technology, infrastructure, and practices to scale our architecture and prepare for future growthCoding or scripting exposureExperience with algorithms, data structures, scripting, pipeline management, and software design

#J-18808-Ljbffr

📌 Manager of Site Reliability Engineering (British Columbia)
🏢 Mastercard
📍 British Columbia

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: manager of site reliability engineering (british columbia) / british columbia

Subscribe to this job alert:

Get the latest job offers by email for: manager of site reliability engineering (british columbia) / british columbia