Who you are
- Experience designing, building, and operating production backend services with strong Rust skills or clear evidence that you can ramp up and contribute in a Rust-first, performance-sensitive codebase
- Experience with distributed-system design, including concurrency, failure handling, consistency, messaging, data partitioning, scalability, and multi-tenant isolation
- Hands-on knowledge of AWS, GCP, or both, including cloud networking, identity and access management, compute, storage, and object storage such as Amazon Simple Storage Service (S3)
- Experience deploying and troubleshooting applications on Kubernetes with Helm and contributing to repeatable, reviewable infrastructure changes using Terraform or similar infrastructure as code tools
- Experience improving the reliability, observability, maintainability, and on-call readiness of backend services, and diagnosing issues across application, data, orchestration, and infrastructure layers to deliver lasting code and system improvements
- Strong system design skills, including making and explaining architectural decisions, documenting constraints, and aligning trade-offs with product and platform needs; ability to work autonomously in ambiguous environments by identifying problems, driving solutions, and taking ownership
- Ability to learn and apply new languages, tools, and frameworks as the problem requires, such as Ruby, Go, or TypeScript and Vue in adjacent parts of the stack, along with excellent written communication and asynchronous collaboration skills demonstrated through thoughtful code review, mentoring, incident communication, and context sharing
- Studies have shown that people from underrepresented groups are less likely to apply to a job unless they meet every single qualification. If you're excited about this role, please apply and allow our recruiters to assess your application
What the job involves
- As a Senior Platform Engineer on the Orbit team, you'll help build and scale GitLab Orbit, a high-impact knowledge graph data service that supports agents, analytics, and architecture-level features across GitLab
- Com, Dedicated, and Self-Managed deployments
- You'll join a small, senior, Rust-first team and focus on how backend code behaves within a distributed, cloud-native system, making graph capabilities reliable, observable, secure, and easy for other teams and agents to use
- In this role, you'll own meaningful parts of the service end to end
- You'll design and implement backend services, data workflows, and interfaces while improving multi-tenant behavior, performance, resilience, and operational readiness
- You'll also contribute to the cloud infrastructure and deployment patterns needed to run the service effectively
- In your first year, you'll take ownership of improving how GitLab Orbit runs in production, with a focus on operational automation, observability, incident readiness, and distributed-system reliability
- You'll also contribute to key areas such as the graph query engine, indexing pipelines, cloud storage integrations, and multi-tenant behavior
- Through thoughtful system design, better tooling, clear runbooks, and shared context, you'll help reduce single points of failure and improve how we build, deploy, and operate the service
- Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment
- Improve the deployment, monitoring,
and operations of GitLab Orbit across GitLab.com, Dedicated, and Self-Managed deployments using Kubernetes, Helm, Terraform, and cloud services from Amazon Web Services, Google Cloud Platform, or both
- Automate recurring operational work and build tools that make deployments, upgrades, recovery, capacity management, and service maintenance safer and more efficient
- Strengthen observability across application, data, orchestration, and infrastructure layers by improving metrics, logs, traces, dashboards, alerts, and service-level indicators to track availability, error rates, and incident response time while collaborating with site reliability engineering teams to improve incident response, on-call readiness, runbooks, and troubleshooting workflows
- Investigate production issues, address underlying causes, and write backend code that handles concurrency, partial failures, retries, consistency, idempotency, performance, and multi-tenant isolation
- Build and improve the graph query engine, software development lifecycle (SDLC) and code indexing pipelines, cloud storage integrations, and application programming interface (API) and Model Context Protocol (MCP) surfaces
- Design reliable, scalable, and cost-aware data workflows using systems such as Amazon Simple Storage Service (S3), ClickHouse, NATS, and Siphon
- Own changes from technical design through rollout and iteration, documenting constraints and trade-offs while collaborating asynchronously with product, data, infrastructure, security, delivery, artificial intelligence, and SRE teams
Advantages
- We offer benefits to manage your health, wealth, and well-being regardless of location
- Flexibility in schedule to be there for life’s important moments
- Equity compensation & Employee Stock Purchase Plan offered
- Generous Paid Time Off
📌 Senior Platform Engineer (Canada)
🏢 GitLab
📍 Canada