03 Oct
|
KEV Group
|
Toronto
KEV builds mission-critical financial software for K12 schools across North America. Our platform provides real-time visibility and control over how funds are collected and managed, replacing fragmented workflows with secure, up-to-date systems that schools rely on every day. KEV builds mission-critical financial software for K12 schools across North America.
Our platform provides real-time visibility and control over how funds are collected and managed, replacing fragmented workflows with secure, modern systems that schools rely on every day. Trusted by more than 27,000 K12 schools managing over $8B annually, KEV delivers mission-critical software for payments, accounting, and reporting, where reliability and security matter. Headquartered in Toronto with teams across North America, we are scaling quickly and investing deeply in cloud-native, data-driven technology.
KEV is a place for people who care about building durable software at scale and doing work that has real impact in public education. We focus on solving hard problems, raising the bar on engineering craft, and building products that earn long-term trust from the communities we serve. We are hiring a Staff Site Reliability Engineer to raise the bar on how reliably KEV's platforms run in production, from the SLOs and error budgets that define what "reliable" means for each service to the incident practice that keeps issues from repeating.
This role requires deep hands-on SRE experience and the judgment to assess where reliability is at risk, shape a strategy for closing the gap, and break that strategy into work a team can execute. You will partner closely with KEV's DevOps Engineering team and product engineering teams to embed reliability into how services are built and operated, and coach and mentor other engineers on reliability practices. Break that strategy into concrete workstreams and lead a team through execution.
Use recurring incidents and operational issues to identify durable fixes and longer-term reliability investments, not just one-off patches. Observability & Monitoring Platform Design and evolve the monitoring, alerting, logging, and tracing platform that gives engineering teams real visibility into system health. Use data to reduce noisy or ineffective alerts and make sure the signals teams rely on are ones they can trust.
Capacity & Performance Engineering Lead capacity planning and performance analysis so that KEV's platforms scale ahead of demand rather than in reaction to it.
Use load testing and performance data to anticipate where systems will break before they do. AI-Assisted Reliability Operations Evaluate and apply AI tooling where it can meaningfully improve reliability and operational workflows — such as anomaly detection or incident triage — applying the same trade‑off thinking to new tooling that you'd apply to any other operational decision.
Coaching & Reliability Mentorship Coach and mentor other engineers on reliability practices — on‑call hygiene, incident response, SLO‑driven prioritization — raising the operational maturity of the broader engineering organization over time. Cross‑Functional Partnership Partner closely with KEV's DevOps Engineering team and product engineering teams to embed reliability into how services are designed, built, and operated. 10+ years of hands‑on experience in Site Reliability Engineering, Platform Engineering, or a similar infrastructure/operations role, with a track record of leading reliability strategy and incident practice at scale. ~ Familiarity working with Microsoft Azure and .NET, including legacy .NET Framework applications and IIS, so reliability practices hold up across KEV's full stack and not just its newest services. ~ Demonstrated experience defining and operating SLOs, SLIs, and error budgets, and using them to guide engineering and operational priorities. ~ Hands‑on experience designing and building observability platforms — monitoring, alerting, logging, and tracing — that give engineering teams real visibility into system health. ~ Strong automation skills, with experience building tools and scripts that eliminate repetitive manual operational work and are treated as production code. ~ Experience with capacity planning and performance engineering, anticipating and preventing reliability issues rather than only reacting to them. ~ Experience evaluating and applying AI tooling to improve reliability and operational workflows. ~ Strong ability to surface trade‑offs and communicate them clearly to both engineering and cross‑functional stakeholders. Competitive compensation – We believe in rewarding great work with fair, competitive pay.
both at work and at home.
Retirement Savings
Support – We help you plan for your future with company‑matched programs, including RRSP matching in Canada and 401(k) contributions in the U.Professional development – We invest in your growth with ongoing learning, stretch opportunities, and continuing education, including KEV Academy for onboarding and skill‑building, plus KEV University, our online platform offering a wide range of courses. Hybrid model – 3 days in the office to collaborate and connect, with flexibility the rest of the week. Flexible PTO – Take the time you need to recharge with close to 4 weeks of vacation and a company‑wide holiday closure Office perks – Enjoy a fully stocked snack bar and occasional catered lunches—because we know that great conversations (and ideas) often start around good food.
Note: We set standard base pay ranges for all roles based on function, level, and country location, benchmarked against similar stage growth companies. Final offer amounts are determined by multiple factors, including skills, depth of work experience and relevant licenses/credentials, and may vary from the amounts listed above. Make a real difference every day – At KEV, your work supports children, parents, and schools across North America.
We don't just build software, we create solutions that simplify lives and strengthen communities. Our mission is rooted in impact, and every team member plays a vital role in shaping the future of K-12 education. How does this help our schools and the students they serve?
Celebrate
Community and Culture – At KEV, we connect, recognize, and celebrate our people. Join Club KEV for team‑building fun, hear directly from customers in our Voice of Customer Series, stay aligned with Monthly Townhalls, and be inspired by the KEVite Awards, where top contributors are recognized by their peers. KEV Group is pleased to accommodate individual needs in accordance with the Accessibility of Ontarians with Disabilities Act, 2005 (AODA), within our recruitment process.
KEV Group is an equal opportunity employer who agrees not to discriminate against any employee or job applicant because of race, color, religion, national origin, sex, physical or mental disability, or age. KEV may use AI‑enabled tools to support components of our recruitment process. These tools are used to support human decision making.
Visit our website for more information and details about working at KEV. #
📌 Site Reliability Engineer, Cloud Networking (Toronto)
🏢 KEV Group
📍 Toronto