06 Sep
|
BigGeo
|
Calgary
1. Company Overview
BigGeo is the Spatial Cloud
We help companies manage and access the world’s spatial data.
Any size, any slice, any insight.
Delivered in seconds.
BigGeo is building the infrastructure layer for spatial intelligence. Our platform enables organizations to store, process, index, query, and analyze massive geospatial datasets at global scale. We are defining the Spatial Cloud category and building the systems that make spatial data accessible, actionable, and performant for modern applications.
2. Why BigGeo Exists and Why People Build HereMost organizations struggle to manage and access spatial data at scale. Data volumes continue to grow, infrastructure becomes increasingly complex, and teams spend more time maintaining systems than generating insights.
BigGeo exists to remove those constraints.
We are building the Spatial Cloud so organizations can work with spatial data of any size, any slice, and derive meaningful insights in seconds rather than hours or days.
Building at BigGeo means working on foundational infrastructure problems that sit at the intersection of cloud computing, distributed systems, geospatial technology, data engineering, and artificial intelligence. Team members are trusted with meaningful ownership, encouraged to think from first principles, and expected to build systems that become core components of a category-defining platform.
We believe AI-native organizations will fundamentally change how companies operate. Every team member is expected to leverage modern AI systems to improve productivity, decision-making, software quality, and operational excellence.
3. Role OverviewBigGeo is seeking a DevOps & Site Reliability Engineer to design, automate, secure, and operate the infrastructure that powers the Spatial Cloud.
This role is responsible for building reliable cloud infrastructure, deployment systems, observability platforms, security controls, and operational tooling that enable engineering teams to deliver production software with confidence.
The role combines DevOps and Site Reliability Engineering responsibilities. You will build the systems that deliver software to production, and you will own the reliability of what runs there — service level objectives, observability, capacity planning, incident response, and the on-call practice that supports them. We combine these deliberately: engineers who build delivery systems make better reliability decisions when they also operate what they ship.
You will work closely with software engineers, platform engineers, data engineers, and product teams to ensure BigGeo’s systems remain scalable, resilient, secure, and highly available as the platform grows.
This role is ideal for someone who enjoys building infrastructure as a product, automating everything possible,
and creating systems that allow engineering teams to move faster without sacrificing reliability.
4.
What You Will
Build and Own
- Cloud infrastructure supporting BigGeo production environments
- Infrastructure-as-Code frameworks and deployment pipelines
- Kubernetes clusters and container orchestration platforms
- CI/CD systems supporting engineering delivery workflows
- Monitoring, logging, alerting, and observability platforms
- Security automation and compliance controls
- Reliability engineering practices and operational standards
- Service level objectives and reliability measurement frameworks
- Incident response processes and on-call practices
- Disaster recovery and business continuity capabilities
- Cost optimization frameworks across cloud environments
- Internal developer platforms and operational tooling
5. Core Responsibilities
- Design, deploy, and maintain scalable cloud infrastructure
- Build and manage Infrastructure-as-Code solutions
- Develop and optimize CI/CD pipelines for engineering teams
- Operate Kubernetes-based production environments
- Improve system reliability, performance, and fault tolerance
- Implement monitoring, observability, and alerting strategies
- Manage cloud networking, security, and access controls
- Automate operational processes and infrastructure workflows
- Lead incident response, root-cause analysis, and blameless postmortems
- Establish reliability standards, SLOs, and operational metrics
- Perform capacity planning for stateful and resource-intensive workloads
- Reduce operational toil through automation and elimination of recurring issues
- Collaborate with engineering teams to improve deployment velocity
- Evaluate and integrate AI-powered operational and automation tools
- Contribute to platform architecture decisions and infrastructure strategy
6. Reliability and On-CallReliability is treated as a core engineering responsibility at BigGeo rather than a separate function. This role carries meaningful ownership of it.
Reliability engineering
- Define and maintain service level objectives for critical services
- Build observability that makes system behavior measurable and actionable
- Perform capacity planning for both growth and failure scenarios
- Drive continuous improvement through postmortems and reliability reviews
On-call
- Production alerting is automated and routed by severity
- On-call responsibility is currently shared across the engineering team and is being formalized into a structured rotation as the team grows
- You will participate in that rotation and help define escalation paths, response expectations, and handoff practices
- New team members shadow incidents before taking primary responsibility
- On-call scheduling and supporting policies are being established as part of this work
7.
Required
Experience
- 4+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering
- Strong experience operating cloud infrastructure in production environments
- Experience building Infrastructure-as-Code solutions
- Experience managing Kubernetes and containerized workloads
- Knowledge of CI/CD pipelines and deployment automation
- Strong Linux systems administration skills
- Experience with monitoring, logging, and observability platforms
- Understanding of networking, security, and distributed systems concepts
- Experience supporting production incident response and troubleshooting
- Experience participating in an on-call rotation for production systems
- Solid communication and cross-functional collaboration skills
8.
Preferred
Experience
- Experience supporting large-scale data platforms
- Experience with geospatial or location-based systems
- Experience building internal developer platforms
- Experience with multi-cloud environments
- Familiarity with high-volume data processing systems
- Experience implementing reliability engineering practices
- Experience defining SLOs, SLIs, and error budgets
- Experience establishing or improving on-call and incident response processes
- Knowledge of database operations and performance tuning
- Experience working within startup or high-growth technology environments
- Experience integrating AI systems into operational workflows
9. Work Environment and CollaborationBigGeo operates as a highly collaborative, AI-native startup environment. This is an on-site role based in Calgary, Alberta. Candidates must be located in the Calgary area or willing to relocate.
Team members are expected to take ownership, move quickly, communicate clearly, and contribute beyond traditional role boundaries when needed. You will work closely with engineering, product, data, and leadership teams while helping establish the operational foundation of a category-defining company.
Success in this role requires curiosity, initiative, systems thinking, and a desire to build infrastructure that enables others to do their best work.
You will have significant influence over how BigGeo scales its platform, operations, and engineering capabilities as we continue defining the Spatial Cloud category.
📌 DevOps and Site Reliability Engineer (Calgary)
🏢 BigGeo
📍 Calgary