26 Sep
|
DuckDuckGo
|
Canada
Who you are
- 10+ years relevant professional experience in reliability, platform, infrastructure, or software engineering, including 4+ years leading SRE teams
- Experience participating in a 24x7 on-call rotation for a large-scale deployment
- Ability to lead and collaborate on high-impact and complex projects from proposal through postmortem
- Proficient in AI-driven development, including designing and implementing agentic workflows
- Skills to wrangle vague problems, propose innovative solutions, and execute them with a strong focus on metrics
- Experience developing effective tools, services, alerts, and responses to identify and address reliability risks
- Investigative ability to root-cause sources of instability in high-traffic, distributed systems
- Deep experience administering and troubleshooting Linux and web technologies
- Ability to implement automation around infrastructure provisioning and configuration management to prioritize efficiency, scalability, and reliability
- Foresight to help identify the future technical direction of our deployment with the goal of improving reliability and performance
- Advanced programming skills enabling close partnership with software engineers to triage production issues and identify appropriate remediation, including code changes and performance considerations
- Ability to leverage cloud-native services and architectures to enhance reliability and scalability, with hands-on experience packaging and deploying applications using Docker and Docker Compose
What the job involves
- Working on the Site Reliability Team,
you'll help build and maintain world-class infrastructure to meet the needs of millions of users protecting their privacy online. You'll utilize high-level languages like Perl, Go, TypeScript, or Python and work on related projects. Recent projects include:
- Ensuring our Duck.ai product meets our reliability standards and minimizing user friction on failures
- Scaling up our own index infrastructure to handle billions of documents
- Create anti fraud verifications that respect users privacy
- As Director, Site Reliability Engineering, you'll dive deep into complex operational challenges, including software, systems, automation, and process analysis. We are looking for candidates who can read, write, troubleshoot, and deploy all types of software to help us tackle the reliability challenges of large-scale deployments
- You’ll be required to attend meetings on camera via video conferencing
- Expect to travel at least two times a year: once for our all-hands meetup and again for a team retreat (each around 4-5 days). While extenuating circumstances may impact attendance, everyone is strongly encouraged to attend
- While we offer a adaptable work arrangement with no core hours, expect an average full-time commitment of 40 hours per week
Benefits
- Remote First, Always:
We've always been a fully distributed company with team members all over the world. We trust you to get your work done wherever, whenever
- Flexible Time Off: We trust you to use good judgment to take time off as needed so that you can be your best at work. Our CEO sets a good example by taking regular vacations and encouraging the rest of the team to do the same
- Office Setup
To help you get your office set up, we reimburse up to $1,250 USD to cover the purchase of office productivity items, such as desk, chair, computer monitors, keyboards, etc
- Co-working: Should you prefer to work from a co-working space, we reimburse the cost up to $500 USD per month
- Wellness Stipend: To promote your physical and mental health, we offer a Wellness Stipend of up to $1,000 USD per year
- Learning Stipend: To promote your professional development and reinforce our culture of self-directed learning and skill-building, we offer a Learning Stipend of up to $1,250 USD per year
- Charitable Donation Matching: We match on qualified charitable donations of up to $1,000 per year
- Assistance Program: We offer counseling services provided through Workplace Options to help you and your family manage life-stressors
- Company Events: We have annual company retreats, team meetups, and optional “work-ations” where you work with colleagues in various destinations around the world. Every quarter we hold 3-day long Hack Days, during which we work on whatever we want
- Healthcare - Medical, Dental, Vision (US Only)
📌 Director of Site Reliability Engineering (Canada)
🏢 DuckDuckGo
📍 Canada