11 Sep
|
NexGen Tech Solutions
|
Canada
11 Sep
NexGen Tech Solutions
Canada
Required qualifications
- 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role.
- Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code.
- Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network and device-integration layers.
- Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments.
- Experience with infrastructure as code and delivery tooling such as Terraform, Helm, Git-based CI/CD and policy-as-code.
- Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis.
- Strong troubleshooting & debugging skills in Kubernetes platforms.
- Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring.
- Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems.
- Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices.
- Experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents.
- Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning.
- Robust networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and systematic packet- or session-level troubleshooting.
- Clear communication, disciplined documentation and the ability to collaborate across NOC, cloud, DevOps, firmware and service-provider teams.
- Bachelor’s degree in computer science, engineering or equivalent practical experience.
Preferred qualifications
- Experience supporting cloud-managed CPEs such as broadband gateways, routers, ONTs, Wi-Fi/mesh systems or similar edge devices in a service-provider environment.
- Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management.
- Experience supporting messaging and streaming platforms such as Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes.
- Understanding of access technologies such as GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless and how CPE, ONTs and provider networks interact.
- Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices.
- Experience building auto-remediation, safe self-service operations or internal reliability platforms.
- Experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment.
- Seniority Level
Mid-Senior level
- Employment Type
Full-time
- Job Functions
- Information Technology
- Skills
- Python (Programming
📌 Site Reliability Engineer (Canada)
🏢 NexGen Tech Solutions
📍 Canada