12 Aug
|
UnivEdge Consulting
|
Canada
12 Aug
UnivEdge Consulting
Canada
Site Reliability Engineer: Level 2
Location: Canada, Remote
Description
We are hiring a Senior Site Reliability Engineer who will contribute in building and supporting reliable, high capacity, and high-performing core infrastructure services that support our mission to reimagine learning for millions of students worldwide.
This position is a full-time contract position. If you have an automation mindset, know AWS services inside out, have distributed system / micro service experience, can triage complex cloud / application stack issues, like to solve engineering problems and lead by example then you will thrive in this position
Qualifications
- Experience as a Software Engineer, Software Developer in Test, SRE, or Devops Engineer with practical experience developing and supporting enterprise applications
- Experience with infrastructure automation technologies like Terraform (3+ years of everyday use preferred)
- Expertise in container/container-fleet-orchestration technologies, specifically ECS and/or EKS
- Versatility with troubleshooting diverse sets of hosting technologies: web server platforms, application platforms, operating systems, network components, virtualization technologies, storage, and database platforms.
- Expertise with continuous-deployment based software development lifecycles (e.g. CI/CD)
- Cloud database operations and deployment experience (RDS MySQL/Postgres/Aurora)
- Expertise with Lean/Agile deployment processes (Blue/Green, ZDT, Canary, DNS migration strategies)
- Familiarity with telemetry SaaS systems like New Relic and Datadog or similar systems from other vendors
- Strong problem solving, triage, root cause analysis and systems engineering skills
- Demonstrated expertise building and managing highly scaled production infrastructure in the cloud
- Experience with AI tooling and integrations, including AI assistants and productivity tools (Copilot, Claude, ChatGPT, etc.).
Specifically looking for
- Experience designing and implementing a multi-region strategy for AWS applications distributed across many AWS accounts that are all using 1 region only currently.
- Experience designing automatic failover, and manual failover mechanisms for multi-region deployed applications such that we are protected automatically from regional failures in addition to having controls to failover manually.
- Experience managing multi-region N-tier applications
- Understanding of what concerns/issues to know when designing multi-region. It’s just as key to know what not to do as it is to know what to do.
Nice to Have MHE is a polyglot organization. Being “conversational” in JavaScript/TypeScript, Python, PHP, Ruby, Golang, Java, Bash, Markdown, reStructuredText, HCL, JSON, YAML, and TOML is a huge plus.
BS Degree in Computer Science (or related technical field and/or equivalent industry experience) preferred
Our cloud stack includes:
20+ application stacks/services spanning multiple programming languages and hosting methods (AWS EKS and ECS). This is just this team, the whole of the company has hundreds of application stacks across hundreds of AWS accounts.
AWS Cloud Tech: AWS ( S3, EC2, ECS, EKS, SQS, Elasticache Redis, Elasticache Memcache, Elasticache Valkey, RDS, Aurora RDS, Open Search/Elastic Search, Dynamodb, Athena, Route53, WAF, IAM, Cloudfront, Load Balancing, ACM, VPCs, Lambda, API Gateway, VPC Endpoints, Routing tables, SES, SNS, Bedrock, Glue and more)
Infrastructure as Code : Terraform
RDBMS: Mysql, PostgreSQL, Aurora including Connection Pooling technologies like Sequelize
Caching: Valkey, Redis, Memcache, Dynamodb
Languages : Python, Nodejs, Java, Golang, PHP, Bash
Container orchestration : AWS ECS, AWS EKS
Telemetry: NewRelic, CloudWatch, Datadog, Splunk
Build/Deploy/Run : Github Actions, Artifactory, GitHub Cloud, Pagerduty, Slack, ServiceNow, Exigence
Security: Rapid7 InsightAppSec, Crowdstrike, Kentik, 1Password, DivvyCloud (Insight Cloudsec), Security Scorecard, WAF, dependabot, checkov, Github Advanced Security scanning, Secret scanning, VDP
Task Management : JIRA
Documentation/Knowledge : Confluence, Office365, DX
📌 Site Reliability Engineer: Level 2 (Canada)
🏢 UnivEdge Consulting
📍 Canada