12 Sep
|
DigiOptimizer
|
Brampton
12 Sep
DigiOptimizer
Brampton
Cloud Platform / Site Reliability Engineer – Kubernetes, Automation & Private Cloud
Department: IT / Operations
Employment Type: Full-time
Location: Remote – Canada
Job Summary
ThinkOn is seeking a Cloud Platform / Site Reliability Engineer to operate, automate, scale, secure, and continuously improve the platforms that power ThinkOn’s private cloud services.
This role combines Site Reliability Engineering, Kubernetes platform engineering, Infrastructure-as-Code, observability, automation, and private cloud operations.
You will support ThinkOn’s current VMware Cloud Foundation (VCF) environments, VMware Kubernetes Service (VKS), GPU-enabled infrastructure and platform services, while helping build capabilities that remain portable to upstream and open-source Kubernetes environments.
The successful candidate will take end-to-end operational ownership of platform services rather than operating within a single technology silo. You will troubleshoot across infrastructure, networking, storage, Kubernetes, observability and platform-service layers; automate repetitive work; improve reliability; and help transition recent services from Architecture into Operations.
Key Responsibilities
Kubernetes & Kubernetes-as-a-Service
Deploy, operate, scale, upgrade, troubleshoot and recover production Kubernetes clusters.
Support VMware Supervisor, VMware Kubernetes Service (VKS), VCF Automation and upstream/open-source Kubernetes environments.
Operate Kubernetes cluster lifecycle technologies including Cluster API and related controllers.
Troubleshoot
Control plane and worker nodes
MachineDeployments and node pools
Controllers and operators
CRDs
Admission webhooks
Scheduling and affinity
RBAC and service accounts
DNS
Networking
Storage
Certificates
Kubernetes APIs
Develop and operate reusable Kubernetes-as-a-Service capabilities for tenant and customer environments.
Support Kubernetes multi-tenancy, namespaces, quotas, policies, isolation and self-service.
Maintain Kubernetes lifecycle processes including provisioning, scaling, remediation, upgrade, backup, recovery and retirement.
Infrastructure-as-Code & Automation
Build and maintain infrastructure using Terraform.
Develop reusable Terraform modules and standardized deployment patterns.
Manage Terraform state, environment promotion and automated change workflows.
Develop Ansible automation for platform and operating-system configuration.
Use Python, Bash, PowerShell, APIs and SDKs to automate platform lifecycle activities.
Integrate platform automation with CI/CD pipelines.
Use vendor-neutral APIs and OpenAPI-based interfaces wherever practical.
Continually identify and eliminate repetitive operational toil.
VMware Cloud Foundation
Deploy, operate and lifecycle-manage VMware Cloud Foundation environments including:
ESXi vCenter vSAN
NSX
SDDC Manager
VCF Operations
VCF Automation
Supervisor
VKS
Perform platform upgrades, patching, lifecycle management and fleet operations.
Support current and future VCF releases while maintaining operational practices that are portable to non-VMware platforms.
Operate VCF through APIs, automation and Infrastructure-as-Code rather than relying solely on manual administration.
Site Reliability Engineering
Apply SRE principles to production cloud and Kubernetes platforms.
Define and maintain meaningful SLIs, SLOs and service-health indicators.
Monitor service availability,
latency, saturation, capacity and error conditions.
Identify recurring failures and engineer permanent solutions rather than repeatedly applying manual fixes.
Participate in production incident response and act as an incident owner where appropriate.
Perform root cause analysis and blameless post-incident reviews.
Track corrective actions through completion.
Develop resilience, failure-recovery and disaster-recovery procedures.
Perform capacity planning and identify platform scaling risks before they become incidents.
Continuously improve MTTR, availability and operational efficiency.
Observability
Build and operate observability capabilities using:
Prometheus
Grafana
OpenTelemetry
Zabbix
Loki
Fluent Bit
Splunk or equivalent technologies
Build Prometheus metrics, exporters, recording rules and alerts.
Develop Grafana dashboards for
VMware infrastructure
GPU infrastructure
Platform services
Capacity
Performance
Application and service health
Integrate infrastructure, Kubernetes and application telemetry to support end-to-end troubleshooting.
Implement distributed metrics, logs, events and telemetry collection using OpenTelemetry and related standards.
Networking & Load Balancing
Troubleshoot networking across infrastructure and Kubernetes layers.
Support
TCP/IP
Routing
DNS
DHCP/IPAM
VLAN/VXLAN
BGP/EVPN
MTU
TLS
Network policy
Load balancing
Operate VMware NSX and VPC-based networking.
Support VMware Avi / NSX Advanced Load Balancer and equivalent load-balancing platforms.
Operate Kubernetes Service, Ingress and Gateway API connectivity.
Support Kubernetes CNI technologies including Antrea, Cilium, Calico, Multus or equivalent.
Support multi-network and multi-NIC Kubernetes architectures where workload, storage or high-performance traffic requires isolation.
Understand Kubernetes IPAM and advanced networking concepts.
Platform Services
Operate, troubleshoot and lifecycle-manage services such as:
Harbor and OCI container registries
HashiCorp Vault and secrets-management systems
Certificate management and PKI
Database-as-a-Service
PostgreSQL and MySQL
Velero and Kubernetes backup/recovery
Redis and Kubernetes operators
S3-compatible object storage
OpenSearch and logging platforms
Kubernetes policy and admission-control services
Developer and platform self-service services
Support installation, configuration, upgrade, migration, backup, recovery, monitoring and operational handover for these services.
Security, Policy & Compliance
Implement and maintain Kubernetes and infrastructure RBAC.
Support federated identity and OIDC-based authentication.
Operate certificate and credential lifecycle processes.
Implement policy-as-code using technologies such as OPA Gatekeeper, Kyverno or equivalent.
Detect and remediate configuration drift.
Support automated security and compliance validation.
Participate in vulnerability and CVE remediation.
Support hardened Kubernetes node images and automated operating-system patching.
Work with Security teams to meet Protected B, ITSG-33, PBMM and other applicable compliance requirements.
Storage & Data Services
Troubleshoot Kubernetes persistent storage, CSI, storage classes, snapshots and volume lifecycle.
Support S3-compatible object storage and object-storage consumption by Kubernetes and application workloads.
Understand data protection requirements across both VM and Kubernetes workloads.
AI & GPU Platform Operations
Support GPU-enabled infrastructure and Kubernetes environments.
Operate and troubleshoot:
NVIDIA GPU Operator vGPU
MIG
GPU scheduling
GPU device plugins
GPU licensing
GPU node configuration
High-performance networking
RDMA where applicable
Platform APIs & Programmability
Use REST APIs, OpenAPI specifications and platform SDKs to automate Day-1 and Day-2 operations.
Develop platform integrations using Python, Terraform, PowerCLI and other appropriate tooling.
Integrate infrastructure and platform APIs into GitLab CI/CD and operational workflows.
Evaluate new platform capabilities and determine how they can be safely automated and operationalized.
Service Transition & Operational Ownership
Work closely with Architecture during the introduction of new services and technologies.
Participate in implementation and progressively assume operational ownership.
Validate operational readiness including:
Monitoring
Alerting
Capacity
Backup
Recovery
Security
Upgrade procedures
Automation
Documentation
Create SOPs, runbooks and troubleshooting guides.
Lead subsequent deployments and upgrades after initial architecture implementation.
Operational Redundancy
Cross-train across critical platform services.
Act as both primary and secondary owner of assigned technologies.
Ensure routine operations and incident-response procedures can be performed by another engineer.
Prevent single-person dependencies through:
Documentation
Automation
Peer review
Knowledge sharing
Pairing
Operational rotation
____
Required Skills & Experience
5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Infrastructure, Kubernetes Engineering or a related field.
Strong production Kubernetes administration and troubleshooting experience.
Strong Linux administration and troubleshooting skills.
Strong hands-on Terraform experience.
Experience with Git-based engineering and Infrastructure-as-Code workflows.
Strong experience with Prometheus and Grafana or comparable observability platforms.
Understanding of Open Telemetry and modern metrics/logging architectures.
Strong networking fundamentals including TCP/IP, routing, DNS, TLS and load balancing.
Experience automating operational work using Python, Bash or equivalent languages.
Experience supporting production incidents and performing root cause analysis.
⸻
Preferred Experience
VMware Cloud Foundation 9.x
VCF Automation
VCF Operations
VKS / VMware Supervisor
Cluster API
Argo CD or Flux
OpenTelemetry
Prometheus / Grafana / Loki
Cilium / Calico / Antrea / Multus
OPA Gatekee
Velero cert-manager
VMware Avi
NSX vSAN
S3-compatible object storage
PostgreSQL or Database-as-a-Service
NVIDIA GPU Operator / vGPU
Kubernetes multi-tenancy
Kubernetes-as-a-Service platforms
Certifications such as CKA, CKS, VMware VCP/VCAP/VCF, Terraform Associate or equivalent are beneficial but not required.
📌 Site Reliability Engineer/cloud Platform (Brampton)
🏢 DigiOptimizer
📍 Brampton