12 Sep
|
DigiOptimizer
|
Brampton
12 Sep
DigiOptimizer
Brampton
Cloud Platform / Site Reliability Engineer – Kubernetes, Automation & Private Cloud
Department: IT / Operations
Employment Type: Full-time
Location: Remote – Canada
Job Summary
ThinkOn is seeking a Cloud Platform / Site Reliability Engineer to operate, automate, scale, secure, and continuously improve the platforms that power ThinkOn’s private cloud services.
This role combines Site Reliability Engineering, Kubernetes platform engineering, Infrastructure-as-Code, observability, automation, and private cloud operations.
You will support ThinkOn’s current VMware Cloud Foundation (VCF) environments, VMware Kubernetes Service (VKS), GPU-enabled infrastructure and platform services, while helping build capabilities that remain portable to upstream and open-source Kubernetes environments.
The successful candidate will take end-to-end operational ownership of platform services rather than operating within a single technology silo. You will troubleshoot across infrastructure, networking, storage, Kubernetes, observability and platform-service layers; automate repetitive work; improve reliability; and help transition new services from Architecture into Operations.
Key Responsibilities
Kubernetes & Kubernetes-as-a-Service
- Deploy, operate, scale, upgrade, troubleshoot and recover production Kubernetes clusters.
- Support VMware Supervisor, VMware Kubernetes Service (VKS), VCF Automation and upstream/open-source Kubernetes environments.
- Operate Kubernetes cluster lifecycle technologies including Cluster API and related controllers.
- Troubleshoot:
- Control plane and worker nodes
- MachineDeployments and node pools
- Controllers and operators
- CRDs
- Admission webhooks
- Scheduling and affinity
- RBAC and service accounts
- DNS
- Networking
- Storage
- Certificates
- Kubernetes APIs
- Develop and operate reusable Kubernetes-as-a-Service capabilities for tenant and customer environments.
- Support Kubernetes multi-tenancy, namespaces, quotas, policies, isolation and self-service.
- Maintain Kubernetes lifecycle processes including provisioning, scaling, remediation, upgrade, backup, recovery and retirement.
Infrastructure-as-Code & Automation
- Build and maintain infrastructure using Terraform.
- Develop reusable Terraform modules and standardized deployment patterns.
- Manage Terraform state, environment promotion and automated change workflows.
- Develop Ansible automation for platform and operating-system configuration.
- Use Python, Bash, PowerShell, APIs and SDKs to automate platform lifecycle activities.
- Integrate platform automation with CI/CD pipelines.
- Use vendor-neutral APIs and OpenAPI-based interfaces wherever practical.
- Continually identify and eliminate repetitive operational toil.
VMware Cloud Foundation
- Deploy, operate and lifecycle-manage VMware Cloud Foundation environments including:
- ESXi
- vCenter
- vSAN
- NSX
- SDDC Manager
- VCF Operations
- VCF Automation
- Supervisor
- VKS
- Perform platform upgrades, patching, lifecycle management and fleet operations.
- Support current and future VCF releases while maintaining operational practices that are portable to non-VMware platforms.
- Operate VCF through APIs, automation and Infrastructure-as-Code rather than relying solely on manual administration.
Site Reliability Engineering
- Apply SRE principles to production cloud and Kubernetes platforms.
- Define and maintain meaningful SLIs, SLOs and service-health indicators.
- Monitor service availability, latency,
saturation, capacity and error conditions.
- Identify recurring failures and engineer permanent solutions rather than repeatedly applying manual fixes.
- Participate in production incident response and act as an incident owner where appropriate.
- Perform root cause analysis and blameless post-incident reviews.
- Track corrective actions through completion.
- Develop resilience, failure-recovery and disaster-recovery procedures.
- Perform capacity planning and identify platform scaling risks before they become incidents.
- Continuously improve MTTR, availability and operational efficiency.
Observability
- Build and operate observability capabilities using:
- Prometheus
- Grafana
- OpenTelemetry
- Zabbix
- Loki
- Fluent Bit
- Splunk or equivalent technologies
- Build Prometheus metrics, exporters, recording rules and alerts.
- Develop Grafana dashboards for:
- VMware infrastructure
- GPU infrastructure
- Platform services
- Capacity
- Performance
- Application and service health
- Integrate infrastructure, Kubernetes and application telemetry to support end-to-end troubleshooting.
- Implement distributed metrics, logs, events and telemetry collection using OpenTelemetry and related standards.
Networking & Load Balancing
- Troubleshoot networking across infrastructure and Kubernetes layers.
- Support:
- TCP/IP
- Routing
- DNS
- DHCP/IPAM
- VLAN/VXLAN
- BGP/EVPN
- MTU
- TLS
- Network policy
- Load balancing
- Operate VMware NSX and VPC-based networking.
- Support VMware Avi / NSX Advanced Load Balancer and equivalent load-balancing platforms.
- Operate Kubernetes Service, Ingress and Gateway API connectivity.
- Support Kubernetes CNI technologies including Antrea, Cilium, Calico, Multus or equivalent.
- Support multi-network and multi-NIC Kubernetes architectures where workload, storage or high-performance traffic requires isolation.
- Understand Kubernetes IPAM and advanced networking concepts.
Platform Services
Operate, troubleshoot and lifecycle-manage services such as:
- Harbor and OCI container registries
- HashiCorp Vault and secrets-management systems
- Certificate management and PKI
- Database-as-a-Service
- PostgreSQL and MySQL
- Velero and Kubernetes backup/recovery
- Redis and Kubernetes operators
- S3-compatible object storage
- OpenSearch and logging platforms
- Kubernetes policy and admission-control services
- Developer and platform self-service services
Support installation, configuration, upgrade, migration, backup, recovery, monitoring and operational handover for these services.
Security, Policy & Compliance
- Implement and maintain Kubernetes and infrastructure RBAC.
- Support federated identity and OIDC-based authentication.
- Operate certificate and credential lifecycle processes.
- Implement policy-as-code using technologies such as OPA Gatekeeper, Kyverno or equivalent.
- Detect and remediate configuration drift.
- Support automated security and compliance validation.
- Participate in vulnerability and CVE remediation.
- Support hardened Kubernetes node images and automated operating-system patching.
- Work with Security teams to meet Protected B, ITSG-33, PBMM and other applicable compliance requirements.
Storage & Data Services
- Troubleshoot Kubernetes persistent storage, CSI, storage classes, snapshots and volume lifecycle.
- Support S3-compatible object storage and object-storage consumption by Kubernetes and application workloads.
- Understand data protection requirements across both VM and Kubernetes workloads.
AI & GPU Platform Operations
- Support GPU-enabled infrastructure and Kubernetes environments.
- Operate and troubleshoot:
- NVIDIA GPU Operator
- vGPU
- MIG
- GPU scheduling
- GPU device plugins
- GPU licensing
- GPU node configuration
- High-performance networking
- RDMA where applicable
Platform APIs & Programmability
- Use REST APIs, OpenAPI specifications and platform SDKs to automate Day-1 and Day-2 operations.
- Develop platform integrations using Python, Terraform, PowerCLI and other appropriate tooling.
- Integrate infrastructure and platform APIs into GitLab CI/CD and operational workflows.
- Evaluate new platform capabilities and determine how they can be safely automated and operationalized.
Service Transition & Operational Ownership
- Work closely with Architecture during the introduction of new services and technologies.
- Participate in implementation and progressively assume operational ownership.
- Validate operational readiness including:
- Monitoring
- Alerting
- Capacity
- Backup
- Recovery
- Security
- Upgrade procedures
- Automation
- Documentation
- Create SOPs, runbooks and troubleshooting guides.
- Lead subsequent deployments and upgrades after initial architecture implementation.
Operational Redundancy
- Cross-train across critical platform services.
- Act as both primary and secondary owner of assigned technologies.
- Ensure routine operations and incident-response procedures can be performed by another engineer.
- Prevent single-person dependencies through:
- Documentation
- Automation
- Peer review
- Knowledge sharing
- Pairing
- Operational rotation
____
Required Skills & Experience
- 5 years of experience in Site Reliability Engineering, Platform Engineering, Cloud Infrastructure, Kubernetes Engineering or a related field.
- Strong production Kubernetes administration and troubleshooting experience.
- Strong Linux administration and troubleshooting skills.
- Strong hands-on Terraform experience.
- Experience with Git-based engineering and Infrastructure-as-Code workflows.
- Robust experience with Prometheus and Grafana or comparable observability platforms.
- Understanding of Open Telemetry and modern metrics/logging architectures.
- Strong networking fundamentals including TCP/IP, routing, DNS, TLS and load balancing.
- Experience automating operational work using Python, Bash or equivalent languages.
- Experience supporting production incidents and performing root cause analysis.
⸻
Preferred Experience
- VMware Cloud Foundation 9.x
- VCF Automation
- VCF Operations
- VKS / VMware Supervisor
- Cluster API
- Argo CD or Flux
- OpenTelemetry
- Prometheus / Grafana / Loki
- Cilium / Calico / Antrea / Multus
- OPA Gatekee
- Velero
- cert-manager
- VMware Avi
- NSX
- vSAN
- S3-compatible object storage
- PostgreSQL or Database-as-a-Service
- NVIDIA GPU Operator / vGPU
- Kubernetes multi-tenancy
- Kubernetes-as-a-Service platforms
Certifications such as CKA, CKS, VMware VCP/VCAP/VCF, Terraform Associate or equivalent are beneficial but not required.
📌 Site Reliability Engineer/cloud Platform (Brampton)
🏢 DigiOptimizer
📍 Brampton