30 Jul
|
Citigroup
|
Mississauga
30 Jul
Citigroup
Mississauga
Overview
A senior-level position responsible for accomplishing results by designing, implementing, and managing the firm's engineering platforms, with a focus on CI/CD, container orchestration, and observability. The objective is to drive automation and streamlining of operations and processes, building and maintaining tools for deployment, monitoring, and operations, and troubleshooting and resolving issues in dev, test, and production environments.
Responsibilities
- Design, build, and maintain the CI/CD infrastructure and tools, with a focus on Tekton and Harness.
- Manage, scale, and secure Open Shift container platforms, ensuring high availability and reliability.
- Develop and manage infrastructure as code (IaC) to automate provisioning and configuration of environments.
- Implement and manage a comprehensive observability stack using tools like Prometheus, Grafana, and others to monitor system health, performance, and reliability.
- Collaborate with development teams to create a seamless developer experience and ensure applications are built with scalability, reliability, and security in mind.
- Utilize in-depth knowledge across multiple infrastructure and development areas to provide technical oversight for the platform.
- Contribute to the formulation of strategies for platform engineering and Dev Ops functional areas.
- Provide evaluative judgment based on the analysis of factual data in complicated and unique situations, including root cause analysis and problem resolution.
- Impact the Dev Ops and Platform Engineering area through monitoring delivery of end results and ensuring essential procedures are followed, while contributing to defining standards.
- Appropriately assess risk when technical decisions are made,
demonstrating consideration for the firm’s reputation and safeguarding clients and assets by driving compliance with applicable laws and policies, and applying ethical judgment.
Leadership Qualities
- Excellent organization skills, robust attention to detail, and the ability to multi-task in a fast-paced environment.
- Demonstrated sense of responsibility and capability to deliver robust solutions quickly and effectively.
- Excellent communication skills, with a requirement to clearly articulate and document technical specifications, processes, and designs.
- Proactive problem-solver with a knack for identifying and resolving issues before they impact production.
- Relationship builder and team player, capable of working with diverse teams across the organization.
- Strong negotiation and prioritization skills to manage competing demands.
- Flexibility to handle multiple complex projects and adapt to changing priorities.
- Promotes teamwork and builds strong relationships within and across global teams.
- Champions continuous process improvement, especially in platform reliability, deployment efficiency, and code quality.
Technical Proficiency
- Hands-on engineer with passion for solving complex infrastructure and automation challenges.
- Strong, hands-on experience with Open Shift administration, configuration, and management.
- Deep expertise in designing and implementing CI/CD pipelines using Tekton and Harness.
- Proven experience in Infrastructure Management and automation using Infrastructure as Code (IaC) tools (e.g., Terraform, Ansible).
- Expertise in building and managing observability tools, including Prometheus for metrics and Grafana for visualization.
- Strong scripting skills (e.g., Python, Bash) for automation and tool development.
- In-depth knowledge of containerization (Docker, Podman) and Kubernetes fundamentals.
- Strong understanding of system design for resiliency, scalability, and performance, backed by robust observability.
- Experience with cloud platforms (AWS, Azure, GCP) is a significant plus.
- Proven ability to lead the implementation of successful, large-scale infrastructure projects.
- Be hands-on with technologies and contribute to architecture, design, and implementation with a focus on quality, scalability, and maintainability.
Qualifications
- 8+ years of relevant experience in Dev Ops, Site Reliability Engineering (SRE), or Platform Engineering.
- Hands-on working experience with container orchestration using Open Shift and Kubernetes.
- Strong, demonstrable experience with CI/CD tools, specifically Tekton and Harness.
- Extensive experience with observability and monitoring stacks, including Prometheus and Grafana.
- Proficiency in Infrastructure as Code (IaC) and configuration management tools.
- Experience with scripting and automation.
- Ability to work proactively and independently to address project requirements, and articulate issues/challenges with enough lead time to mitigate project delivery risks.
- A history of conducting code reviews and ensuring high standards for infrastructure and automation code.
- Basic knowledge of industry practices and standards in the Dev Ops and SRE space.
- Consistently demonstrates clear and concise written and verbal communication.
Education
- Bachelor’s degree/University degree or equivalent experience.
- Master’s degree is a plus.
#J-18808-Ljbffr
📌 Apps Development Sr Manager - Vice President (Mississauga)
🏢 Citigroup
📍 Mississauga