Site Reliability Engineer/cloud Platform (Brampton)

Site Reliability Engineer/cloud Platform (Brampton)

12 Sep
|
DigiOptimizer
|
Brampton

12 Sep

DigiOptimizer

Brampton

Cloud Platform / Site Reliability Engineer – Kubernetes, Automation & Private Cloud

Department: IT / Operations

Employment Type: Full-time

Location: Remote – Canada

Job Summary

ThinkOn is seeking a Cloud Platform / Site Reliability Engineer to operate, automate, scale, secure, and continuously improve the platforms that power ThinkOn’s private cloud services.

This role combines Site Reliability Engineering, Kubernetes platform engineering, Infrastructure-as-Code, observability, automation, and private cloud operations.

You will support ThinkOn’s current VMware Cloud Foundation (VCF) environments, VMware Kubernetes Service (VKS), GPU-enabled infrastructure and platform services, while helping build capabilities that remain portable to upstream and open-source Kubernetes environments.

The successful candidate will take end-to-end operational ownership of platform services rather than operating within a single technology silo. You will troubleshoot across infrastructure, networking, storage, Kubernetes, observability and platform-service layers; automate repetitive work; improve reliability; and help transition new services from Architecture into Operations.

Key Responsibilities

Kubernetes & Kubernetes-as-a-Service

- Deploy, operate, scale, upgrade, troubleshoot and recover production Kubernetes clusters.
- Support VMware Supervisor, VMware Kubernetes Service (VKS), VCF Automation and upstream/open-source Kubernetes environments.
- Operate Kubernetes cluster lifecycle technologies including Cluster API and related controllers.
- Troubleshoot:
- Control plane and worker nodes
- MachineDeployments and node pools
- Controllers and operators
- CRDs
- Admission webhooks
- Scheduling and affinity
- RBAC and service accounts
- DNS
- Networking
- Storage
- Certificates
- Kubernetes APIs
- Develop and operate reusable Kubernetes-as-a-Service capabilities for tenant and customer environments.
- Support Kubernetes multi-tenancy, namespaces, quotas, policies, isolation and self-service.
- Maintain Kubernetes lifecycle processes including provisioning, scaling, remediation, upgrade, backup, recovery and retirement.

Infrastructure-as-Code & Automation

- Build and maintain infrastructure using Terraform.
- Develop reusable Terraform modules and standardized deployment patterns.
- Manage Terraform state, environment promotion and automated change workflows.
- Develop Ansible automation for platform and operating-system configuration.
- Use Python, Bash, PowerShell, APIs and SDKs to automate platform lifecycle activities.
- Integrate platform automation with CI/CD pipelines.
- Use vendor-neutral APIs and OpenAPI-based interfaces wherever practical.
- Continually identify and eliminate repetitive operational toil.

VMware Cloud Foundation

- Deploy, operate and lifecycle-manage VMware Cloud Foundation environments including:
- ESXi
- vCenter
- vSAN
- NSX
- SDDC Manager
- VCF Operations
- VCF Automation
- Supervisor
- VKS
- Perform platform upgrades, patching, lifecycle management and fleet operations.
- Support current and future VCF releases while maintaining operational practices that are portable to non-VMware platforms.
- Operate VCF through APIs, automation and Infrastructure-as-Code rather than relying solely on manual administration.

Site Reliability Engineering

- Apply SRE principles to production cloud and Kubernetes platforms.
- Define and maintain meaningful SLIs, SLOs and service-health indicators.
- Monitor service availability, latency,



saturation, capacity and error conditions.
- Identify recurring failures and engineer permanent solutions rather than repeatedly applying manual fixes.
- Participate in production incident response and act as an incident owner where appropriate.
- Perform root cause analysis and blameless post-incident reviews.
- Track corrective actions through completion.
- Develop resilience, failure-recovery and disaster-recovery procedures.
- Perform capacity planning and identify platform scaling risks before they become incidents.
- Continuously improve MTTR, availability and operational efficiency.

Observability

- Build and operate observability capabilities using:
- Prometheus
- Grafana
- OpenTelemetry
- Zabbix
- Loki
- Fluent Bit
- Splunk or equivalent technologies
- Build Prometheus metrics, exporters, recording rules and alerts.
- Develop Grafana dashboards for:
- VMware infrastructure
- GPU infrastructure
- Platform services
- Capacity
- Performance
- Application and service health
- Integrate infrastructure, Kubernetes and application telemetry to support end-to-end troubleshooting.
- Implement distributed metrics, logs, events and telemetry collection using OpenTelemetry and related standards.

Networking & Load Balancing

- Troubleshoot networking across infrastructure and Kubernetes layers.
- Support:
- TCP/IP
- Routing
- DNS
- DHCP/IPAM
- VLAN/VXLAN
- BGP/EVPN
- MTU
- TLS
- Network policy
- Load balancing
- Operate VMware NSX and VPC-based networking.
- Support VMware Avi / NSX Advanced Load Balancer and equivalent load-balancing platforms.
- Operate Kubernetes Service, Ingress and Gateway API connectivity.
- Support Kubernetes CNI technologies including Antrea, Cilium, Calico, Multus or equivalent.
- Support multi-network and multi-NIC Kubernetes architectures where workload, storage or high-performance traffic requires isolation.
- Understand Kubernetes IPAM and advanced networking concepts.

Platform Services

Operate, troubleshoot and lifecycle-manage services such as:

- Harbor and OCI container registries
- HashiCorp Vault and secrets-management systems
- Certificate management and PKI
- Database-as-a-Service
- PostgreSQL and MySQL
- Velero and Kubernetes backup/recovery
- Redis and Kubernetes operators
- S3-compatible object storage
- OpenSearch and logging platforms
- Kubernetes policy and admission-control services
- Developer and platform self-service services

Support installation, configuration, upgrade, migration, backup, recovery, monitoring and operational handover for these services.

Security, Policy & Compliance

- Implement and maintain Kubernetes and infrastructure RBAC.
- Support federated identity and OIDC-based authentication.
- Operate certificate and credential lifecycle processes.
- Implement policy-as-code using technologies such as OPA Gatekeeper, Kyverno or equivalent.
- Detect and remediate configuration drift.
- Support automated security and compliance validation.
- Participate in vulnerability and CVE remediation.
- Support hardened Kubernetes node images and automated operating-system patching.




- Work with Security teams to meet Protected B, ITSG-33, PBMM and other applicable compliance requirements.

Storage & Data Services

- Troubleshoot Kubernetes persistent storage, CSI, storage classes, snapshots and volume lifecycle.
- Support S3-compatible object storage and object-storage consumption by Kubernetes and application workloads.
- Understand data protection requirements across both VM and Kubernetes workloads.

AI & GPU Platform Operations

- Support GPU-enabled infrastructure and Kubernetes environments.
- Operate and troubleshoot:
- NVIDIA GPU Operator
- vGPU
- MIG
- GPU scheduling
- GPU device plugins
- GPU licensing
- GPU node configuration
- High-performance networking
- RDMA where applicable

Platform APIs & Programmability

- Use REST APIs, OpenAPI specifications and platform SDKs to automate Day-1 and Day-2 operations.
- Develop platform integrations using Python, Terraform, PowerCLI and other appropriate tooling.
- Integrate infrastructure and platform APIs into GitLab CI/CD and operational workflows.
- Evaluate new platform capabilities and determine how they can be safely automated and operationalized.

Service Transition & Operational Ownership

- Work closely with Architecture during the introduction of new services and technologies.
- Participate in implementation and progressively assume operational ownership.
- Validate operational readiness including:
- Monitoring
- Alerting
- Capacity
- Backup
- Recovery
- Security
- Upgrade procedures
- Automation
- Documentation
- Create SOPs, runbooks and troubleshooting guides.
- Lead subsequent deployments and upgrades after initial architecture implementation.

Operational Redundancy

- Cross-train across critical platform services.
- Act as both primary and secondary owner of assigned technologies.
- Ensure routine operations and incident-response procedures can be performed by another engineer.
- Prevent single-person dependencies through:
- Documentation
- Automation
- Peer review
- Knowledge sharing
- Pairing
- Operational rotation

____

Required Skills & Experience

- 5 years of experience in Site Reliability Engineering, Platform Engineering, Cloud Infrastructure, Kubernetes Engineering or a related field.
- Strong production Kubernetes administration and troubleshooting experience.
- Strong Linux administration and troubleshooting skills.
- Strong hands-on Terraform experience.
- Experience with Git-based engineering and Infrastructure-as-Code workflows.
- Robust experience with Prometheus and Grafana or comparable observability platforms.
- Understanding of Open Telemetry and modern metrics/logging architectures.
- Strong networking fundamentals including TCP/IP, routing, DNS, TLS and load balancing.
- Experience automating operational work using Python, Bash or equivalent languages.
- Experience supporting production incidents and performing root cause analysis.



Preferred Experience

- VMware Cloud Foundation 9.x
- VCF Automation
- VCF Operations
- VKS / VMware Supervisor
- Cluster API
- Argo CD or Flux
- OpenTelemetry
- Prometheus / Grafana / Loki
- Cilium / Calico / Antrea / Multus
- OPA Gatekee
- Velero
- cert-manager
- VMware Avi
- NSX
- vSAN
- S3-compatible object storage
- PostgreSQL or Database-as-a-Service
- NVIDIA GPU Operator / vGPU
- Kubernetes multi-tenancy
- Kubernetes-as-a-Service platforms

Certifications such as CKA, CKS, VMware VCP/VCAP/VCF, Terraform Associate or equivalent are beneficial but not required.

📌 Site Reliability Engineer/cloud Platform (Brampton)
🏢 DigiOptimizer
📍 Brampton

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer/cloud platform (brampton) / brampton

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer/cloud platform (brampton) / brampton