The hardest question in this work is not whether it works on the day you ship. It is whether it is still working a month later, when the model provider has changed something, the client's data has drifted, and nobody has looked at a trace in weeks.
This role owns that question. You will build the evaluation and observability layer that turns our reliability claims into something we can actually prove, to ourselves and to clients.
What you will do
- Build evaluation harnesses and golden datasets for client systems, including the subjective criteria that need an LLM judge and the calibration to trust one
- Own tracing from prompt to output: every retrieval, tool call and result captured well enough to debug an incident from
- Detect regressions before clients do, across model upgrades, prompt changes, and data drift
- Design guardrails and human-in-the-loop checkpoints for the decisions that are too consequential to automate outright
- Turn recurring failure modes into standards the whole delivery team works to
What we look for
- Strong Python and a genuinely empirical instinct: you reach for a measurement before an opinion
- Experience evaluating LLM systems, including where offline metrics stop predicting production behaviour
- Observability and monitoring background, whether from ML or classical distributed systems
- Statistical literacy sufficient to know when a difference between two runs means nothing
- Willingness to be the person who says a system is not ready
Nice to have
- You have built an LLM-as-judge evaluator and validated it against human labels
- Regulated or safety-critical environment experience
- Data engineering skills for the pipelines behind evaluation sets
Your first six months
- Month 1: instrument a live engagement from prompt to output and find something nobody knew was broken
- Month 3: every active engagement has an evaluation suite running against it
- Month 6: regressions are caught by our tooling rather than by a client email