14 Sep
|
BigGeo
|
Alberta
DevOps and Site Reliability Engineer
About BigGeo
BigGeo is the Spatial Cloud. We help companies manage and access the world’s spatial data. Any size, any slice, any insight. Delivered in seconds.
We’re building something that hasn’t existed before: a new layer of the internet where the “where” and “when” behind every decision is instantly clear, programmable, and actionable. Our platform removes the complexity that has kept spatial data locked in silos for decades, and replaces it with speed, precision, and control.
We’re a Calgary-based company, early and moving fast, with real customers, real infrastructure, and a clear point of view on where the world is going.
Why BigGeo Exits and Why People Build Here
Most companies are spatially blind. They know what their data says, but not where or when things actually happen. That gap costs real money, creates real risk, and limits what AI can actually do in the physical world. BigGeo exists to close that gap. We’re not building another tool. We’re building the rails that connect the planet’s moving data to the systems that run the world. That’s a big problem, and it takes people who care about doing things right, not just fast.
People build here because:
The problem is real and the category is open. We’re not competing for the middle of an existing market. We’re defining a new one. Your work shapes what the category becomes.
Your fingerprints are on the architecture. We’re at the stage where the decisions you make today become the foundation tomorrow. What you ship matters.
We run on clarity, not politics. We move with purpose. No bureaucratic drag, just a team that agrees on the mission and gets to work.
You’ll grow fast because the problems are hard. Spatial data at scale is a genuinely difficult domain. If you want to be stretched, you’ll be stretched.
We’re building for longevity. We’re not chasing hype cycles. We’re building infrastructure, the kind that compounds in value over time and earns the trust of the companies that depend on it.
The Role
BigGeo is seeking a DevOps & Site Reliability Engineer to design, automate, secure, and operate the infrastructure that powers the Spatial Cloud.
This role is responsible for building reliable cloud infrastructure, deployment systems, observability platforms, security controls, and operational tooling that enable engineering teams to deliver production software with confidence.
The role combines DevOps and Site Reliability Engineering responsibilities. You will build the systems that deliver software to production, and you will own the reliability of what runs there — service level objectives, observability, capacity planning, incident response, and the on-call practice that supports them. We combine these deliberately: engineers who build delivery systems make better reliability decisions when they also operate what they ship.
You will work closely with software engineers, platform engineers, data engineers, and product teams to ensure BigGeo’s systems remain scalable, resilient, secure, and highly available as the platform grows.
This role is ideal for someone who enjoys building infrastructure as a product, automating everything possible, and creating systems that allow engineering teams to move faster without sacrificing reliability.
What You Will Build and Own
Cloud infrastructure supporting BigGeo production environments
Infrastructure-as-Code frameworks and deployment pipelines
Kubernetes clusters and container orchestration platforms
CI/CD systems supporting engineering delivery workflows
Monitoring, logging, alerting, and observability platforms
Security automation and compliance controls
Reliability engineering practices and operational standards
Service level objectives and reliability measurement frameworks
Incident response processes and on-call practices
Disaster recovery and business continuity capabilities
Cost optimization frameworks across cloud environments
Internal developer platforms and operational tooling
Key Responsibilities
Core Responsibilities
Design, deploy, and maintain scalable cloud infrastructure
Build and manage Infrastructure-as-Code solutions
Develop and optimize CI/CD pipelines for engineering teams
Operate Kubernetes-based production environments
Improve system reliability, performance, and fault tolerance
Implement monitoring, observability, and alerting strategies
Manage cloud networking, security, and access controls
Automate operational processes and infrastructure workflows
Lead incident response, root-cause analysis, and blameless postmortems
Establish reliability standards, SLOs, and operational metrics
Perform capacity planning for stateful and resource-intensive workloads
Reduce operational toil through automation and elimination of recurring issues
Collaborate with engineering teams to improve deployment velocity
Evaluate and integrate AI-powered operational and automation tools
Contribute to platform architecture decisions and infrastructure strategy
Reliability and On-Call
Reliability is treated as a core engineering responsibility at BigGeo rather than a separate function. This role carries meaningful ownership of it.
Reliability engineering
Define and maintain service level objectives for critical services
Build observability that makes system behavior measurable and actionable
Perform capacity planning for both growth and failure scenarios
Drive continuous improvement through postmortems and reliability reviews
On-call
Production alerting is automated and routed by severity
On-call responsibility is currently shared across the engineering team and is being formalized into a structured rotation as the team grows
You will participate in that rotation and help define escalation paths, response expectations, and handoff practices
New team members shadow incidents before taking primary responsibility
On-call scheduling and supporting policies are being established as part of this work
Technology and Tools
The exact stack will evolve as the platform grows, but experience with many of the following technologies is expected:
Cloud Platforms
AWS, Azure, or Google Cloud Platform
Infrastructure and Automation
Terraform
Terragrunt
Pulumi
Infrastructure-as-Code frameworks
GitOps workflows
Containers and Orchestration
Docker
Kubernetes
Helm
CI/CD
GitHub Actions
GitLab CI
ArgoCD
Flux
Jenkins
Observability
Prometheus
Grafana
Datadog
OpenTelemetry
ELK/OpenSearch
Reliability and Incident Management
SLO and SLI frameworks
Incident management and paging tools
Status and escalation workflows
Security
IAM
Secrets management
Cloud security tooling
Vulnerability management
Collaboration and Operations
Slack
Google Workspace
Monday
Linear
AI Tooling
Modern AI assistants and development tools
AI-supported operational automation
AI-enhanced monitoring and troubleshooting workflows
What You Bring
Required Experience
4+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering
Strong experience operating cloud infrastructure in production environments
Experience building Infrastructure-as-Code solutions
Experience managing Kubernetes and containerized workloads
Knowledge of CI/CD pipelines and deployment automation
Strong Linux systems administration skills
Experience with monitoring, logging, and observability platforms
Understanding of networking, security, and distributed systems concepts
Experience supporting production incident response and troubleshooting
Experience participating in an on-call rotation for production systems
Robust communication and cross-functional collaboration skills
Preferred Experience
Experience supporting large-scale data platforms
Experience with geospatial or location-based systems
Experience building internal developer platforms
Experience with multi-cloud environments
Familiarity with high-volume data processing systems
Experience implementing reliability engineering practices
Experience defining SLOs, SLIs, and error budgets
Experience establishing or improving on-call and incident response processes
Knowledge of database operations and performance tuning
Experience working within startup or high-growth technology environments
Experience integrating AI systems into operational workflows
How This Role Contributes to the Spatial Cloud
The Spatial Cloud depends on infrastructure that is reliable, scalable, secure, and globally available.
As a DevOps & Site Reliability Engineer, you will build and operate the foundational systems that allow BigGeo to manage and access the world’s spatial data at massive scale. Your work directly impacts platform reliability, developer productivity, customer experience, and the company’s ability to deliver spatial insights in seconds.
The systems you build become critical components of the infrastructure powering the next generation of spatial computing.
Work Environment and Collaboration
BigGeo operates as a highly collaborative, AI-native startup environment.
This is an on-site role based in Calgary, Alberta. Candidates must be located in the Calgary area or willing to relocate.
Team members are expected to take ownership, move quickly, communicate clearly, and contribute beyond traditional role boundaries when needed. You will work closely with engineering, product, data, and leadership teams while helping establish the operational foundation of a category-defining company.
Success in this role requires curiosity, initiative, systems thinking, and a desire to build infrastructure that enables others to do their best work.
You will have significant influence over how BigGeo scales its platform, operations, and engineering capabilities as we continue defining the Spatial Cloud category.
Success Metrics
Success in this role means quickly developing a deep understanding of BigGeo’s production setting, taking meaningful ownership of platform reliability, and improving the systems that allow engineering teams to deploy and operate the Spatial Cloud safely at scale.
First 30 Days — Understand the Environment and Establish Operational Context
By the end of your first 30 days, you will:
Develop a working understanding of BigGeo’s cloud infrastructure, production architecture, Kubernetes environments, networking, deployment systems, and Infrastructure-as-Code.
Understand the current CI/CD workflows and how software moves from development through deployment into production.
Review existing monitoring, logging, alerting, and observability coverage for critical services.
Understand current production reliability expectations, incident response practices, escalation paths, and the developing on-call model.
Shadow production incidents and participate in troubleshooting alongside engineering team members before taking primary on-call responsibility.
Identify initial reliability, security, infrastructure, deployment, or operational risks and document recommended priorities.
Establish working relationships with software, platform, data, and product teams and understand the infrastructure requirements of their workloads.
Begin using AI-assisted tools for infrastructure analysis, troubleshooting, automation, documentation, and operational workflows.
Success at 30 days: You can explain how BigGeo’s production infrastructure operates, how software reaches production, how system health is measured, where the most important operational risks exist, and how you will begin improving reliability and developer operations.
First 60 Days — Take Ownership and Improve Reliability
By the end of your first 60 days, you will:
Independently own meaningful components of BigGeo’s cloud infrastructure, Kubernetes environments, deployment systems, or observability stack.
Deliver at least one material improvement to Infrastructure-as-Code, CI/CD, deployment automation, observability, or operational tooling.
Establish or materially improve SLOs and SLIs for critical production services, creating clearer visibility into reliability and service health.
Improve monitoring and alerting so critical production conditions are actionable while unnecessary or low-value alerts are reduced.
Participate actively in the on-call rotation with appropriate escalation support and demonstrate effective production troubleshooting and incident response.
Lead or materially contribute to root-cause analysis and blameless postmortems, ensuring recurring issues result in clear corrective actions.
Identify recurring operational toil and automate at least one high-value manual infrastructure or operational workflow.
Assess capacity requirements for critical stateful or resource-intensive workloads and identify scaling or failure risks.
Strengthen infrastructure security, access controls, secrets management, or vulnerability management where gaps are identified.
Use AI-assisted operational workflows to accelerate troubleshooting, infrastructure development, analysis, or automation without compromising reliability or security.
Success at 60 days: You are independently operating important parts of BigGeo’s infrastructure, contributing effectively to production response, and delivering measurable improvements to reliability, automation, observability, and engineering delivery.
First 90 Days — Strengthen the Operational Foundation of the Spatial Cloud
By the end of your first 90 days, you will:
Own the reliability and operational health of defined production infrastructure and services, with clear SLOs, monitoring, alerting, and response expectations.
Demonstrate measurable improvement in at least one key operational area such as deployment reliability, deployment velocity, incident frequency, recovery time, alert quality, infrastructure automation, or operational toil.
Help establish a structured on-call practice with clear escalation paths, response expectations, handoff practices, and supporting documentation.
Establish repeatable incident response and postmortem practices that convert production failures into system and process improvements.
Improve infrastructure resilience through capacity planning, fault-tolerance improvements, recovery planning, or elimination of identified single points of failure.
Advance disaster recovery and business continuity capabilities for critical infrastructure and workloads.
Deliver reusable infrastructure automation or internal developer tooling that allows engineering teams to deploy and operate services with less manual intervention.
Identify cloud cost or resource-efficiency opportunities and implement practical improvements without compromising reliability or performance.
Establish clear operational documentation and runbooks for critical infrastructure, deployment, incident, and recovery workflows.
Contribute informed recommendations to platform architecture and infrastructure strategy based on direct production experience.
Demonstrate that infrastructure improvements are creating leverage across the engineering organization by making production systems easier to deploy, observe, troubleshoot, recover, and scale.
Success at 90 days: You are operating as a trusted owner of BigGeo’s production reliability. Engineering teams can ship with greater confidence, critical systems are more observable and resilient, operational practices are becoming repeatable, and infrastructure is better prepared to support the continued scale of the Spatial Cloud.
#J-18808-Ljbffr
📌 DevOps and Site Reliability Engineer (Alberta)
🏢 BigGeo
📍 Alberta