Senior Site Reliability Engineer/DevOps (Toronto)

Senior Site Reliability Engineer/DevOps (Toronto)

03 Oct
|
TATA Consultancy Services
|
Toronto

03 Oct

TATA Consultancy Services

Toronto

Tata Consultancy Services (TCS) is an equal opportunity employer, and embraces diversity in race, nationality, ethnicity, gender, age, physical ability, neurodiversity, and sexual orientation, to create a workforce that reflects the societies we operate in. Our continued commitment to Culture and Diversity is reflected in our people stories across our workforce and implemented through equitable workplace policies and processes. TCS has been serving the Canadian marketplace for more than 35 years and today is the technology partner of choice for many of the country's leading organizations.

As one of the largest IT services providers in Canada and globally, TCS is aiming to become the world's largest AI-led technology services company and is enabling its clients to transform themselves across the full AI stack, from infrastructure to intelligence. Rooted in the heritage and values of the Tata Group, TCS is focused on creating long‑term value for its clients, investors, employees, and the communities it serves. With a highly skilled workforce supporting customers from coast to coast, TCS has consistently been recognized as a Top Employer by the Top Employers Institute, including a recent top ranking at the country level.

TCS also sponsors 14 of the world's most prestigious marathons and endurance events, including the TCS Toronto Waterfront Marathon, reflecting its commitment to health, sustainability, and community empowerment. As an Intermediate Site Reliability Engineer, you will support and continuously improve enterprise Azure and Databricks platforms. The role focuses on production reliability, monitoring, incident response, availability, operational readiness and platform support.

You will work with platform engineering, security, network, application and data teams to keep services stable, secure and supportable. Monitor and support production Azure and Databricks environments, ensuring availability,



performance and operational readiness.

Troubleshoot

Azure platform, Databricks, networking, storage, identity, access and application‑related issues.

Support

Databricks workspaces, compute, cluster policies, jobs, workflows, user access, monitoring and cost controls. Support integrations between Databricks, Azure Data Lake Storage Gen2, Azure Data Factory, Azure SQL, Key Vault and managed identities.

Support

Azure networking and connectivity components, including VNets, NSGs, routes, private endpoints, DNS, VPN/ExpressRoute and hub‑and‑spoke connectivity.

Support Azure

Storage services, including storage accounts, Blob Storage and ADLS Gen2, with appropriate access, availability and lifecycle controls. Manage alerts, dashboards and operational monitoring using Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog or New Relic. Contribute to root cause analysis, problem management and follow‑up remediation for recurring incidents.

Assist with patching, upgrades, maintenance windows, service validation and disaster recovery exercises. Manage work through JIRA and ServiceNow and contribute to daily standups, planning sessions and service reviews. Work with engineering teams to improve reliability, reduce recurring issues and strengthen operational support.

Positive years of hands‑on experience supporting Databricks environments.

Experience working with Azure networking concepts, including VNets, NSGs, private endpoints, DNS and routing.



Hands‑on experience with monitoring and observability platforms such as Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog or New Relic.

Experience following incident escalation, change management, root cause analysis and problem management processes.

Experience working with JIRA, ServiceNow, operational runbooks and enterprise support procedures. Good years of Windows Server administration experience. Good years of Linux administration experience. Basic understanding of network troubleshooting and connectivity concepts.

Experience supporting Databricks jobs, clusters, workflows, workspaces and user access. Good years of Azure SQL operational support experience. Good years of Azure Data Factory operational support experience, including linked services, integration runtimes and Databricks orchestration.

Experience supporting AI/GenAI platforms, Azure OpenAI, model endpoints, RAG services or MLOps operations.

Experience supporting enterprise data and analytics platforms.

Experience with capacity review, platform health reporting, cost monitoring and performance troubleshooting. Strong collaboration skills and the ability to work with engineering, platform, network, security and support teams. Strong organizational skills and the ability to manage work through JIRA and ServiceNow.

Customer‑focused mindset with a commitment to reliability and operational excellence. Ability to explain technical issues clearly and maintain accurate documentation. Bachelor's degree in Computer Science, Engineering, Information Technology or a related field, or equivalent practical experience.

Tata Consultancy Services Canada

Inc. is committed to meeting the accessibility needs of all individuals in accordance with the Accessibility for Ontarians with Disabilities Act (AODA) and the Ontario Human Rights Code (OHRC).

📌 Senior Site Reliability Engineer/DevOps (Toronto)
🏢 TATA Consultancy Services
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer/devops (toronto) / toronto