Evals Lead, AI Observability
LOCATION: Remote | Type: Contract
About Newpage Solutions
Newpage Solutions is a global digital health innovation company helping people live longer, healthier lives. We partner with life sciences organisations which include, pharmaceutical, biotech and healthcare leaders, to build transformative AI and data driven technologies addressing real-world health challenges.
From strategy and research to UX design and agile development, we deliver and validate impactful solutions using lean, human-centered practices.
We are proud to be a ‘Great Place to Work®’ certified company for the last three consecutive years. We also hold a top Glassdoor rating and are named among the "Top 50 Most Promising Healthcare Solution Providers" by CIOReview. As an organisation, we foster creativity, continuous learning and inclusivity, creating an workplace where bold ideas thrive and make a measurable difference in people’s lives.
Your Mission
We are looking for an evaluation lead to help project teams evaluate their LLM-based applications / agents well, using a shared observability and evaluation platform (Langfuse). This is the person who makes that platform useful to the teams adopting it, and who brings the evaluation expertise they do not have themselves. Building and running the platform is handled by separate engineering roles, and is not part of this one.
Accountability for any individual product's evaluation results stays with the team that owns it. This role is not a release gate and does not run other teams' evaluations for them.
What You’ll Do
Enablement and adoption content.
- Explain how to set up an evaluation. Written guidance that takes a team from a use case to something running: deciding what to measure, choosing a method that suits it,
assembling a first dataset, and knowing when a result can be trusted.
- Set out what good practice looks like, and why. Building and labelling a golden dataset, holding out test data, choosing a sample size, writing and calibrating a judge
Write evaluation patterns and templates.
- Produce reusable patterns for the evaluation problems that come up across projects, covering task success, faithfulness and grounding, retrieval quality, tool and action correctness, safety and refusal behaviour, and regression against known failures.
- Maintain templates, worked examples and reference implementations, built on the platform's own capabilities, that a team can copy and adapt rather than work out from first principles.
Provide consultancy to projects on evaluation design
- Work with individual teams on how to evaluate their particular use case, including which dimensions matter and how to measure them credibly.
- Review evaluation designs, datasets and judge prompts, and recommend thresholds.
- Advise on measurement soundness: sample size, variance, and whether a reported improvement is supported by the data.
- Leave accountability for the product's evaluation results with the team that owns it. The role advises rather than approves.
What You Bring
- Direct experience designing and running LLM evaluations, not only reading the results of someone else's.
The candidate should have designed suites, built datasets and made decisions about measurement, and should be able to describe an evaluation they built that changed a shipping decision.
- LLM-as-judge in practice, including its calibration against human judgement and its limitations.
- Dataset construction and labelling. Sampling strategy, annotation guidelines, measuring agreement between annotators, and managing contamination and drift over time.
- Sound grasp of measurement. Sample size and statistical significance, variance between runs, and the difference between a real improvement and noise.
- Working knowledge of LLM application architecture, including retrieval, tool calling, agent loops and prompt management
- Strong written and verbal communication
- Good to have Experience with LLM observability and evaluation tooling such as Langfuse, LangSmith, Braintrust, Arize Phoenix, DeepEval or Ragas.
What We Offer
At Newpage, we’re building a company that works smart and grows with agility, where driven individuals come together to do work that matters. We offer:
- A people-first culture - Supportive peers, open communication and a strong sense of belonging
- Smart, purposeful collaboration - Work with talented colleagues to create technologies that solve meaningful business challenges
- Balance that lasts - We respect your time and support a healthy integration of work and life
- Room to grow - Opportunities for learning, leadership and career development, shaped around you
- Meaningful rewards - Competitive compensation that recognises both contribution and potential
Ready to Apply?
Let’s build the future of health together. Apply below or reach out to:
[email protected]
📌 Evals Lead, AI Observability (Canada)
🏢 NewPage Solutions
📍 Canada