07 Oct
|
Thinkon
|
Etobicoke
ThinkOn is seeking a Cloud Platform / Site Reliability Engineer to operate, automate, monitor, and improve our private cloud and Kubernetes platforms.
The role combines hands-on Kubernetes operations, Infrastructure-as-Code, observability, automation, Linux systems administration, and VMware Cloud Foundation (VCF) operations. You will support ThinkOn’s VCF and VMware Kubernetes Service (VKS) environments and contribute to Kubernetes and Kubernetes-as-a-Service (KaaS) capabilities.
You will work closely with Architecture, Infrastructure, Networking, Security, and Operations teams and progressively assume operational ownership of platform services.
We are looking for engineers who do more than resolve individual incidents. When something breaks, the goal is to detect troubleshoot restore understand the cause prevent recurrence improve monitoring document and share the solution. Engineers should look for opportunities to make systems more automated, observable, repeatable, and resilient.
ThinkOn is a remote-first organization. At this time, we are welcoming candidates from Canada for this position. Please note that this listing is for a current vacancy at ThinkOn.
It is a mandatory requirement for the position that the candidate must be eligible to obtain Canadian Secret Clearance. To meet this eligibility, the candidate must be a Canadian citizen or permanent resident for at least 10 consecutive years and must have at least 10 years of verifiable background information. This is essential to ensure compliance with federal security standards and to support the sensitive nature of the projects involved. Candidates who do not meet this criteria will not be considered for the role.
You Will:
- Kubernetes & Platform Operations
- Deploy, operate, troubleshoot, scale, upgrade, back up, recover, and decommission Kubernetes clusters, including VMware Supervisor and VKS.
- Troubleshoot workloads, worker nodes, controllers, operators, CRDs, scheduling, storage, networking, RBAC, DNS, and certificates.
- Support multi-tenant Kubernetes using namespaces, quotas, RBAC, and policy controls, and help develop Kubernetes-as-a-Service capabilities.
- Terraform, Helm & Automation
- Build and maintain infrastructure with Terraform, using reusable Infrastructure-as-Code modules and configurations.
- Deploy and maintain Kubernetes platform services with Helm; troubleshoot values, templates, dependencies, upgrades, and rollbacks.
- Work with Git-based engineering workflows and CI/CD pipelines.
- Identify repetitive operational tasks and automate them using Ansible, Bash, Python, PowerShell, or APIs.
- VMware Cloud Foundation
- Support VCF environments: vSphere / ESXi / vCenter, vSAN, NSX, SDDC Manager, VCF Operations, VCF Automation, Supervisor, and VMware Kubernetes Service.
- Assist with platform upgrades, patching, capacity expansion, and lifecycle management, using APIs and automation rather than manual administration wherever possible.
- Site Reliability, Monitoring & Observability
- Apply SRE principles to improve reliability and efficiency: monitor availability, performance, capacity, errors, and system health, and contribute to SLOs, capacity planning, and recovery procedures.
- Participate in incident response, root cause analysis, and post-incident reviews, driving long-term corrective actions for recurring issues.
- Build and maintain observability using Prometheus, Grafana, Zabbix, OpenTelemetry, Loki, Fluent Bit, and Splunk.
- Create dashboards, alerts,
Prometheus queries, metrics, and exporters; tune alerts to reduce noise and integrate monitoring with ticketing and incident-management workflows.
- Networking, Storage & Load Balancing
- Troubleshoot networking issues involving TCP/IP, DNS, routing, TLS, MTU, load balancing, and network policies.
- Support VMware NSX, VMware Avi / NSX Advanced Load Balancer, and Kubernetes networking (Services, Ingress, Gateway API, CNI).
- Troubleshoot Kubernetes storage — StorageClasses, persistent volumes, snapshots, and CSI integrations — and support vSAN and other platform storage services.
- Platform Services
- Install, configure, monitor, upgrade, back up, recover, and troubleshoot platform services such as Harbor and container registries, HashiCorp Vault and secrets management, certificate management and PKI, PostgreSQL / MySQL and Database-as-a-Service, Velero, Redis and Kubernetes operators, S3-compatible object storage, and logging and observability platforms.
- GPU & AI Infrastructure
- Support GPU-enabled infrastructure and Kubernetes workloads, including NVIDIA GPU Operator, vGPU, GPU scheduling, device plugins, and GPU monitoring with Prometheus/Grafana.
- Security & Compliance
- Support Kubernetes and infrastructure RBAC, federated identity and OIDC integrations, certificates, secrets, and secure-access controls.
- Assist with platform hardening, vulnerability remediation, configuration management, multi-tenant isolation, and Zero Trust principles.
- Work with Security and Compliance teams to support Protected B, ITSG-33, PBMM, ISO 27001, SOC 2, and related requirements.
- Service Transition, Ownership & Redundancy
- Work with Architecture to introduce new platform services, including implementation, testing, and operational readiness, validating monitoring, alerting, backup, recovery, and support procedures before production handover.
- Create and maintain SOPs, runbooks, and troubleshooting documentation, and share knowledge with the wider Operations team.
- Progressively assume ownership of recurring platform operations; cross-train with other platform engineers and provide backup coverage to reduce single-person dependencies.
- Participate in shared operational ownership and on-call responsibilities where required.
You Have:
- Eligible to obtain required Canadian Federal Security Clearance upon your first day of employment is a hard requirement
- 3–5+ years of experience in Platform Engineering, Systems Engineering, SRE, Cloud Infrastructure, Kubernetes, or a related role.
- Hands-on production Kubernetes experience.
- Strong Linux administration and troubleshooting skills.
- Hands-on experience with Terraform.
- Hands-on experience with Helm.
- Experience with Git and basic CI/CD workflows.
- Experience with Prometheus and Grafana or similar monitoring technologies.Good understanding of networking fundamentals including
- TCP/IP, DNS, TLS, routing, and load balancing.
- Understanding of Kubernetes networking, storage, RBAC, scheduling, and workload lifecycle.
- Experience scripting or automating tasks using Bash, Python,
PowerShell, or similar.
- Strong troubleshooting and diagnostic skills.
- Experience documenting technical procedures and operational processes. Ability to learn new technologies and work across both commercial and open-source platforms.
- Ability to collaborate effectively with Architecture, Infrastructure, Networking, Security, and Operations teams.
Experience with some of the following is beneficial but not required:
- VMware Cloud Foundation 9.x, VCF Automation / VCF Operations, VMware Supervisor / VKS, vSAN, VMware Avi / NSX
- Cluster API, Argo CD / Flux, GitLab CI/CD, Ansible, OpenTelemetry
- Cilium / Calico / Antrea / Multus, OPA Gatekeeper / Kyverno, cert-manager
- Harbor, HashiCorp Vault, Velero, S3-compatible object storage, PostgreSQL / DBaaS
- NVIDIA GPU Operator / vGPU, Kubernetes-as-a-Service environments, government or regulated environments
- Relevant Kubernetes, Terraform, Linux, VMware, or cloud certifications
Perks and Perks:
- A remote-first culture
- Competitive compensation package
- Flexible time-off for vacation; 5 personal days/year
- Comprehensive health and dental benefits
- A $500* Wellness & Learning allowance
- Eligible to join the ThinkOn Group Retirement Savings Plan
- Referral Bonuses for successful hires
About ThinkOn Inc. | Where Data Thrives:
Founded in 2013, ThinkOn is a managed infrastructure services provider (MISP) with a global data center footprint, focused on empowering partners to do whatever they need to do with their data. ThinkOn’s team of data-obsessed experts protect clients’ data like it’s their own, making it more resilient, secure, actionable, and searchable. ThinkOn is Channel First and works with a global network of value-add resellers and managed service providers to provide creative, turnkey Infrastructure-as-a-Service (IaaS), Disaster Recovery-as-a-Service (DRaaS), and Backup-as-a-Service (BaaS) solutions and data management services that are fast, flexible, scalable, highly secure, and cost-effective with predictable pricing and no hidden fees. ThinkOn has data centers located across North America, the United Kingdom, Australia, and the Caribbean.
Recognized for its substantial growth and global success, ThinkOn has been named to several notable lists, including The Globe and Mail’s Report on Business ranking of “Canada’s Top Growing Companies” (2021 and 2022), the Deloitte Technology Fast 500 (2021 and 2022), the Deloitte Technology Fast 50 (2021), CIOReview’s “Most Promising Backup Solution Provider” (2021), the Canadian Business “Growth List,” and Channel Daily News’ “Top 100 Solution Providers” (2021).
ThinkOn is headquartered in Toronto, Ontario. To learn more, visit www.ThinkOn.com
Accessibility Accommodations:
ThinkOn is committed to a workforce that is reflective of diverse populations. We welcome applications from qualified individuals from all backgrounds. In accordance with the Accessibility for Ontarians with Disabilities Act (AODA) and accessibility standards across Canada, ThinkOn provides accommodations to job applicants with disabilities throughout the recruitment process. If you require accommodations, please let us know and we will work with you to meet your needs. We are committed to a selection process and work environment that is inclusive, equitable, accessible, and adheres to our corporate values.
Learn more about our corporate values at www.ThinkOn.com/values/.
📌 Cloud Platform & Site Reliability Engineer (Etobicoke)
🏢 Thinkon
📍 Etobicoke