03 Aug
|
Pythian
|
Ottawa
At Pythian, we are experts in strategic database and analytics services, driving digital transformation and operational excellence. Pythian, a multinational company, was founded in 1997 and started by ensuring the reliability and performance of mission‑critical databases. We quickly earned a reputation for solving tough data challenges. We were there when the industry moved from on‑premises to cloud environments, and as enterprises sought more from their data, we expanded our competencies to include advanced analytics.
Today, we empower organizations to embrace transformation and leverage advanced technologies, including AI, to stay competitive. We deliver innovative solutions that meet each client’s data goals and have built strong partnerships with Google Cloud, AWS, Microsoft, Oracle, SAP, and Snowflake. The powerful combination of our extensive expertise in data and cloud and our ability to keep on top of the latest bleeding‑edge technologies makes us the perfect partner to help mid and large‑sized businesses transform to stay ahead in today’s rapidly changing digital economy.
Why you
Pythian is building a next‑generation Site Reliability Engineering team, and we’re looking for talented, motivated engineers who thrive in fast‑paced, problem‑solving environments. As an SRE, you’ll design, deploy, and operate large‑scale distributed systems across compute, storage, networking, and AI/ML environments. You’ll lead projects from architecture to automation to intelligent monitoring, collaborating with both clients and teammates to build resilient, high‑performing infrastructure.
What you will be doing
- Operate and optimize Kubernetes clusters, Istio service mesh, and Linux‑based systems.
- Automate workflows using Go,
Python, and Shell scripting.
- Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
- Troubleshoot complex networking, storage, and system performance issues.
- Partner with AI/ML teams to ensure infrastructure readiness for model training and data pipelines.
- Participate in on‑call rotations and postmortem reviews to improve system resilience.
What we need from you
- Experience with Google Cloud, plus IaC tools (Terraform).
- Strong knowledge of microservices, containers (Kubernetes, Docker), and networking.
- Hands‑on experience with PKI, service mesh, and Linux systems administration.
- SRE mindset with a focus on automation, scalability, and reliability.
- Experience with Golang is an asset.
What you will receive
- Love your career – Competitive total rewards package. Work during hours you choose; take a day off to volunteer for your favorite charity.
- Love your coworkers – Collaborate with some of the best and brightest in the industry.
- Love your development – Hone your skills or learn new ones with our substantial training allowance; participate in professional development days, attend training, become certified, whatever you like.
- Love your workspace – We give you all the equipment you need to work from home including a laptop with your choice of OS, and an annual budget to personalize your work setting.
- Love yourself – Pythian cares about the health and well‑being of our team. You will have an annual wellness budget to make yourself a priority (gym memberships, massages, fitness and more). Additionally, you will receive generous paid vacation and sick days, as well as a day off to volunteer for your favorite charity.
#J-18808-Ljbffr
📌 Site Reliability Consultant (Ottawa)
🏢 Pythian
📍 Ottawa