Senior AI Evals Engineer (AI Quality & Safety)
Libra
Seniority
Senior
Model
Hybrid
Sector
Salary
Undisclosed
Contract
Full-Time
As an Evals Engineer at Libra, you'll be the first dedicated hire on a new team that owns how we measure quality: the datasets, rubrics, judges and harnesses that decide whether an AI feature is good enough to ship, and the automated optimization that runs against them.
What you'll do
- Build and own the eval platform: datasets, judges, harnesses, regression gates, cost and quality in one view, on Python, FastAPI and Langfuse.
- Map the quality landscape: good means something different for research, drafting, summarisation and retrieval, and again per jurisdiction.
- Design the rubrics and set the standard for LLM-as-judge: turn Legal Engineers' and subject-matter experts' judgment into version-controlled criteria.
- Deep-dive results and traces until you can say why something failed, then find the lever that moves it and automate the fix.
- Unhobble the optimizer: instrument the app so prompts, hyperparameters and harness architecture become levers a search can safely pull.
- Make cheaper models win: treat quality per euro as a first-class metric, instrumented per call.
- Own guardrails: ungrounded advice, invented or misattributed citations, jurisdiction and language leakage, prompt injection from ingested documents.
What you'll need
- Bachelor's degree or equivalent in a relevant technical field (e.g. Computer Science, Software Engineering, Statistics, Data Science).
- Minimum 5 years in software engineering, at least 1 building LLM-powered products in production.
- Strong Python: FastAPI, modern tooling, and the data stack (pandas, numpy, notebooks).
- Hands-on experience designing evaluations: datasets, rubrics, LLM-as-judge, benchmarking, human labelling.
- Solid security and data-privacy practice.
- AI coding agents (Claude Code, Codex, Cursor) in your daily workflow.
- Genuine interest in the legal domain.
- Excellent communication in English.
Nice to have
- Automated prompt or pipeline optimization (GEPA, DSPy or similar).
- Broader data science toolkit, e.g. embedding clustering to check dataset coverage.

