06 Aug
|
Tern Travel
|
Dover
Tern's user base is about to triple. Large host agencies are coming on board this year, and the infrastructure needs to be ready before they arrive. If you want to own the migration, build the monitoring, and be the person every other engineer at Tern depends on, this is your role.
Everything Tern ships rides on the infrastructure underneath it. We're a Ruby on Rails application on Heroku, migrating to Google Cloud Platform, with a Postgres core and a data pipeline through Fivetran, BigQuery, and Hex. It's a solid foundation. It won't scale on its own.
We're preparing for a major step‑up in load, large host agencies coming on board, user base tripling, and this role builds the infrastructure that holds under that growth. You'll own the migration to GCP, own the monitoring and alerting that keeps production reliable, and own the hot paths that need to get faster before volume climbs. Do it well and every other engineer at Tern moves faster and sleeps better. This is a force multiplier role at the foundation of the product.
This is a player‑coach role. You'll start as the hands‑on technical lead for infrastructure and grow into coaching and managing alongside it. That's the expectation from day one.
Our stack: Tern is a Ruby on Rails application with a Hotwire front end, backed by a Postgres database and hosted on Heroku, though we are migrating to Google Cloud Platform. Our data flows through to BigQuery, where we build reporting in Hex. Claude Code, with our own library of custom skills and agents, is part of daily development. You don't need to have used every piece, but you should be fluent enough to be productive quickly and excited to work this way.
WHAT YOU'LL DO
- Own the migration from Heroku to Google Cloud Platform: architecture, execution, and a cutover that doesn't surprise anyone
- Build and maintain the Postgres core, Fivetran pipeline: BigQuery data layer, and Hex reporting infrastructure
- Optimize the hot paths that matter most: key backend code paths and our heaviest third party syncs, so performance holds as volume climbs
- Own monitoring, alerting, cost reduction, and proactive scaling: surface problems early, keep spend sane, and stay ahead of growth rather than reacting to it
- Lead incident response and write post‑mortems that turn an outage into a permanent fix and a smarter team
- Set the operational bar across engineering and pull others up to it
WHAT YOU HAVE
- Production reliability ownership: Track record of personally owning production reliability at meaningful scale. Concrete stories of incidents you led, fixed, and prevented from recurring, not just participated in. This is a primary responsibility, not something you've done on the side.
- Infrastructure migrations: Real experience owning a cloud migration end to end, not just contributing to one. Fluent in GCP (or a comparable cloud), infrastructure‑as‑code, and the failure modes of distributed systems.
- Observability and proactive operations: You build monitoring and alerting that surfaces problems before users find them. You know what to instrument, what to alert on, and what's just noise.
- High agency: You find the highest‑leverage reliability problem and go fix it without being assigned to it. You don't wait for an outage to justify the work.
- AI in your working habits: Specific examples of how AI has made your debugging, automation, or operational workflows faster or more reliable.
BONUS POINTS
- GCP migration experience, specifically from Heroku or another PaaS
- Experience with Fivetran, BigQuery, or Hex in a production data pipeline
- Has managed or coached infrastructure engineers
WHY JOIN TERN?
- Be part of a mission‑driven team transforming the travel planning space
- Work with a supportive, curious, and creative team
- Own the infrastructure layer of a product used by thousands of travel advisors running their businesses
- Market-competitive salary, equity, and benefits package
📌 Site Reliability Engineer (Dover)
🏢 Tern Travel
📍 Dover