03 Oct
|
rentsync_careers
|
Montreal
03 Oct
rentsync_careers
Montreal
Rentsync is an award-winning, high-growth organization that provides high quality websites, marketing services, and software solutions to the rental and property management industry throughout Canada and the United States. We're looking for a hands-on Site Reliability Engineer to lead our response to production incidents. You'll dig into our AWS and Kubernetes environments, find the root cause, and fix it. Between incidents, you'll make sure the same problem doesn't happen twice by hardening infrastructure, improving monitoring, and working with engineering teams on performance and reliability. Our workplace is bigger and more varied than most companies our size: 10+ products and 100+ services across multiple Kubernetes clusters, mainly on AWS with some Azure and GCP, built in PHP, Ruby on Rails, JavaScript/TypeScript, .NET, Python, and Rust. AWS (EKS, EC2, RDS, S3, ALB/NLB, CloudWatch), Azure, GCP, Kubernetes, Terraform, Ansible, GitHub Actions/GitLab CI, PagerDuty, Prometheus/Mimir, Loki, Tempo, Grafana, OpenTelemetry, Cloudflare, Ubuntu &
• Amazon Linux, MySQL &
• PostgreSQL, Redis &
• Memcached, NGINX &
• Traefik, Bash, Python, and applications built in PHP, Ruby on Rails, JavaScript/TypeScript, .NET, and Rust. This is a remote position . Diagnose and fix issues directly in AWS (EKS, EC2, RDS, networking, IAM) and Kubernetes, such as failing pods, resource exhaustion, bad deploys, networking/DNS, database and cache problems. Run blameless post-mortems and personally drive the technical follow-up work, not just the action-item list. Automate runbooks and repetitive operational work, including using AI tools to speed up triage, investigation, and remediation. Reliability engineering (preventing the next incident) Build and maintain monitoring for Kubernetes workloads and services (Prometheus/Mimir, Loki, Tempo,
Grafana, OpenTelemetry), with low-noise, high-signal alerts.
Create and maintain production test suites: synthetic checks, smoke tests, health checks, and load/performance tests. Partner with engineering teams to find and fix performance and reliability issues, and define SLOs, SLIs, and error budgets. Keep service docs and architecture decisions current so any engineer can operate our systems. Infrastructure as code with Terraform, and comfort working in CI/CD pipelines.
Solid
Linux, networking, and container fundamentals. Scripting/automation in Bash, Python, or similar. 3+ years in a cloud engineering, DevOps, or SRE role supporting production web applications. ~ Strong, hands-on AWS experience in production (EKS, EC2, RDS, VPC networking, IAM, CloudWatch). ~ Deep production Kubernetes experience, including troubleshooting, debugging, and monitoring Kubernetes workloads. ~ A track record of working with engineering teams to identify and resolve performance and reliability issues. ~ Experience with monitoring and observability tools (e.g.
Experience building automated tests or checks for production reliability (synthetics, smoke, health, or load testing). ~ Willingness to take part in an after-hours on-call rotation as we introduce one in the future. Using AI tools to accelerate SRE work, such as incident triage, log and metric analysis, runbook automation, or infrastructure code. Supporting many tech stacks across multiple teams (PHP, Ruby on Rails, .NET, Python, Rust, JavaScript). AWS certification (e.g.
Solutions
Architect, DevOps Engineer, or SysOps). Load testing tools such as k6, Locust, or JMeter. MySQL/PostgreSQL operations and Redis/Memcached tuning. Cloud cost optimization and capacity planning. Rentsync reserves the right to use Artificial Intelligence to screen and/or assess candidates. #
📌 Senior Site Reliability Engineer/DevOps (Montreal)
🏢 rentsync_careers
📍 Montreal