30 Jul
|
HCLTech
|
Canada
Toronto, Ontario
Job Summary
The Technical Support Specialist in Site Reliability engineering (SRE) will be responsible for ensuring the reliability and stability of the systems and applications. The role involves providing technical support, troubleshooting issues, and implementing solutions to optimize system performance and availability. The Specialist will play a crucial role in monitoring, analyzing, and enhancing system reliability to meet business objectives effectively.
Key Responsibilities
1. Provide technical support and assistance to ensure the reliability and stability of systems and applications
2. Troubleshoot and resolve technical issues related to system performance and availability
3. Collaborate with cross functional teams to implement and optimize solutions for system reliability
4. Monitor system performance metrics and develop strategies to enhance system reliability
5. Implement automation tools and scripts to improve system monitoring and maintenance processes
Skill Requirements
1. Strong knowledge of site reliability engineering (sre) principles and practices
2. Proficiency in scripting languages like python, shell, or powershell
3. Experience with monitoring tools such as prometheus, grafana, or nagios
4. Familiarity with cloud platforms like aws, azure, or google cloud
5. Excellent problem-solving and analytical skills
6. Strong communication and teamwork abilities
7.
Ability to work in a fast paced and dynamic environment
Other Requirements
Core Skills & Tools
AWS: Lambda, ECS/Fargate/EC2, API Gateway, SNS/SQS, Kinesis, RDS; IAM/KMS foundations.
Observability & ITSM: Dynatrace, CloudWatch, ELK; ServiceNow for incidents/changes; SLI/SLO dashboards.
Reliability Practices: Error budgets, capacity/performance benchmarking, automation/runbook execution, FinOps awareness.
1.Relevant certifications in Site Reliability Engineering (SRE) or related fields are a plus
Deliver 24×7 monitoring, incident response, and problem management; drive MTTA/MTTR reduction and SLO/SLI adherence.
Perform preventive health checks; analyze ticket trends to implement continual service improvements and automation to reduce toil.
Execute blameless postmortems and high-quality RCA; maintain SOPs/runbooks and reliability dashboards.
Configure/tune observability (Dynatrace, CloudWatch, ELK); enable self-healing workflows and workload optimizations.
Support change/service requests within agreed SLAs; collaborate during transitions and onboard current AWS services.
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-
📌 SRE Technical Specialist (Canada)
🏢 HCLTech
📍 Canada