Site Reliability Engineer/cloud Platform (Brampton)

Site Reliability Engineer/cloud Platform (Brampton)

12 Sep
|
DigiOptimizer
|
Brampton

12 Sep

DigiOptimizer

Brampton

Cloud Platform / Site Reliability Engineer – Kubernetes, Automation & Private Cloud

Department: IT / Operations

Employment Type: Full-time

Location: Remote – Canada

Job Summary

ThinkOn is seeking a Cloud Platform / Site Reliability Engineer to operate, automate, scale, secure, and continuously improve the platforms that power ThinkOn’s private cloud services.

This role combines Site Reliability Engineering, Kubernetes platform engineering, Infrastructure-as-Code, observability, automation, and private cloud operations.

You will support ThinkOn’s current VMware Cloud Foundation (VCF) environments, VMware Kubernetes Service (VKS), GPU-enabled infrastructure and platform services, while helping build capabilities that remain portable to upstream and open-source Kubernetes environments.

The successful candidate will take end-to-end operational ownership of platform services rather than operating within a single technology silo. You will troubleshoot across infrastructure, networking, storage, Kubernetes, observability and platform-service layers; automate repetitive work; improve reliability; and help transition recent services from Architecture into Operations.

Key Responsibilities

Kubernetes & Kubernetes-as-a-Service

Deploy, operate, scale, upgrade, troubleshoot and recover production Kubernetes clusters.

Support VMware Supervisor, VMware Kubernetes Service (VKS), VCF Automation and upstream/open-source Kubernetes environments.

Operate Kubernetes cluster lifecycle technologies including Cluster API and related controllers.

Troubleshoot

Control plane and worker nodes

MachineDeployments and node pools

Controllers and operators

CRDs

Admission webhooks

Scheduling and affinity

RBAC and service accounts

DNS

Networking

Storage

Certificates

Kubernetes APIs

Develop and operate reusable Kubernetes-as-a-Service capabilities for tenant and customer environments.

Support Kubernetes multi-tenancy, namespaces, quotas, policies, isolation and self-service.

Maintain Kubernetes lifecycle processes including provisioning, scaling, remediation, upgrade, backup, recovery and retirement.

Infrastructure-as-Code & Automation

Build and maintain infrastructure using Terraform.

Develop reusable Terraform modules and standardized deployment patterns.

Manage Terraform state, environment promotion and automated change workflows.

Develop Ansible automation for platform and operating-system configuration.

Use Python, Bash, PowerShell, APIs and SDKs to automate platform lifecycle activities.

Integrate platform automation with CI/CD pipelines.

Use vendor-neutral APIs and OpenAPI-based interfaces wherever practical.

Continually identify and eliminate repetitive operational toil.

VMware Cloud Foundation

Deploy, operate and lifecycle-manage VMware Cloud Foundation environments including:

ESXi vCenter vSAN

NSX

SDDC Manager

VCF Operations

VCF Automation

Supervisor

VKS

Perform platform upgrades, patching, lifecycle management and fleet operations.

Support current and future VCF releases while maintaining operational practices that are portable to non-VMware platforms.

Operate VCF through APIs, automation and Infrastructure-as-Code rather than relying solely on manual administration.

Site Reliability Engineering

Apply SRE principles to production cloud and Kubernetes platforms.

Define and maintain meaningful SLIs, SLOs and service-health indicators.

Monitor service availability,



latency, saturation, capacity and error conditions.

Identify recurring failures and engineer permanent solutions rather than repeatedly applying manual fixes.

Participate in production incident response and act as an incident owner where appropriate.

Perform root cause analysis and blameless post-incident reviews.

Track corrective actions through completion.

Develop resilience, failure-recovery and disaster-recovery procedures.

Perform capacity planning and identify platform scaling risks before they become incidents.

Continuously improve MTTR, availability and operational efficiency.

Observability

Build and operate observability capabilities using:

Prometheus

Grafana

OpenTelemetry

Zabbix

Loki

Fluent Bit

Splunk or equivalent technologies

Build Prometheus metrics, exporters, recording rules and alerts.

Develop Grafana dashboards for

VMware infrastructure

GPU infrastructure

Platform services

Capacity

Performance

Application and service health

Integrate infrastructure, Kubernetes and application telemetry to support end-to-end troubleshooting.

Implement distributed metrics, logs, events and telemetry collection using OpenTelemetry and related standards.

Networking & Load Balancing

Troubleshoot networking across infrastructure and Kubernetes layers.

Support

TCP/IP

Routing

DNS

DHCP/IPAM

VLAN/VXLAN

BGP/EVPN

MTU

TLS

Network policy

Load balancing

Operate VMware NSX and VPC-based networking.

Support VMware Avi / NSX Advanced Load Balancer and equivalent load-balancing platforms.

Operate Kubernetes Service, Ingress and Gateway API connectivity.

Support Kubernetes CNI technologies including Antrea, Cilium, Calico, Multus or equivalent.

Support multi-network and multi-NIC Kubernetes architectures where workload, storage or high-performance traffic requires isolation.

Understand Kubernetes IPAM and advanced networking concepts.

Platform Services

Operate, troubleshoot and lifecycle-manage services such as:

Harbor and OCI container registries

HashiCorp Vault and secrets-management systems

Certificate management and PKI

Database-as-a-Service

PostgreSQL and MySQL

Velero and Kubernetes backup/recovery

Redis and Kubernetes operators

S3-compatible object storage

OpenSearch and logging platforms

Kubernetes policy and admission-control services

Developer and platform self-service services

Support installation, configuration, upgrade, migration, backup, recovery, monitoring and operational handover for these services.

Security, Policy & Compliance

Implement and maintain Kubernetes and infrastructure RBAC.

Support federated identity and OIDC-based authentication.

Operate certificate and credential lifecycle processes.

Implement policy-as-code using technologies such as OPA Gatekeeper, Kyverno or equivalent.

Detect and remediate configuration drift.

Support automated security and compliance validation.

Participate in vulnerability and CVE remediation.

Support hardened Kubernetes node images and automated operating-system patching.





Work with Security teams to meet Protected B, ITSG-33, PBMM and other applicable compliance requirements.

Storage & Data Services

Troubleshoot Kubernetes persistent storage, CSI, storage classes, snapshots and volume lifecycle.

Support S3-compatible object storage and object-storage consumption by Kubernetes and application workloads.

Understand data protection requirements across both VM and Kubernetes workloads.

AI & GPU Platform Operations

Support GPU-enabled infrastructure and Kubernetes environments.

Operate and troubleshoot:

NVIDIA GPU Operator vGPU

MIG

GPU scheduling

GPU device plugins

GPU licensing

GPU node configuration

High-performance networking

RDMA where applicable

Platform APIs & Programmability

Use REST APIs, OpenAPI specifications and platform SDKs to automate Day-1 and Day-2 operations.

Develop platform integrations using Python, Terraform, PowerCLI and other appropriate tooling.

Integrate infrastructure and platform APIs into GitLab CI/CD and operational workflows.

Evaluate new platform capabilities and determine how they can be safely automated and operationalized.

Service Transition & Operational Ownership

Work closely with Architecture during the introduction of new services and technologies.

Participate in implementation and progressively assume operational ownership.

Validate operational readiness including:

Monitoring

Alerting

Capacity

Backup

Recovery

Security

Upgrade procedures

Automation

Documentation

Create SOPs, runbooks and troubleshooting guides.

Lead subsequent deployments and upgrades after initial architecture implementation.

Operational Redundancy

Cross-train across critical platform services.

Act as both primary and secondary owner of assigned technologies.

Ensure routine operations and incident-response procedures can be performed by another engineer.

Prevent single-person dependencies through:

Documentation

Automation

Peer review

Knowledge sharing

Pairing

Operational rotation

____

Required Skills & Experience

5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Infrastructure, Kubernetes Engineering or a related field.

Strong production Kubernetes administration and troubleshooting experience.

Strong Linux administration and troubleshooting skills.

Strong hands-on Terraform experience.

Experience with Git-based engineering and Infrastructure-as-Code workflows.

Strong experience with Prometheus and Grafana or comparable observability platforms.

Understanding of Open Telemetry and modern metrics/logging architectures.

Strong networking fundamentals including TCP/IP, routing, DNS, TLS and load balancing.

Experience automating operational work using Python, Bash or equivalent languages.

Experience supporting production incidents and performing root cause analysis.



Preferred Experience

VMware Cloud Foundation 9.x

VCF Automation

VCF Operations

VKS / VMware Supervisor

Cluster API

Argo CD or Flux

OpenTelemetry

Prometheus / Grafana / Loki

Cilium / Calico / Antrea / Multus

OPA Gatekee

Velero cert-manager

VMware Avi

NSX vSAN

S3-compatible object storage

PostgreSQL or Database-as-a-Service

NVIDIA GPU Operator / vGPU

Kubernetes multi-tenancy

Kubernetes-as-a-Service platforms

Certifications such as CKA, CKS, VMware VCP/VCAP/VCF, Terraform Associate or equivalent are beneficial but not required.

📌 Site Reliability Engineer/cloud Platform (Brampton)
🏢 DigiOptimizer
📍 Brampton

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer/cloud platform (brampton) / brampton

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer/cloud platform (brampton) / brampton