Guide · 2026
LLM evaluation, without the folklore
How to actually measure whether a language model or agent is doing its job: offline versus online evaluation, the metrics worth tracking, why most agent failures hide at the conversation level, and the tools that run it.
The short version
- Assert everything you can for free. Valid JSON, schema, required fields, latency, cost. No judge needed.
- Use a model judge only for what you can't assert. Groundedness, tone, task completion.
- Evaluate the conversation, not just the span. Agent failures live in the trajectory.
- Then check the judge. An uncalibrated judge is a confident random number generator.
What LLM evaluation actually is
Evaluation is scoring output against criteria you defined, so quality becomes something you measure instead of something you sense. Normal software testing asserts equality — this function returns 4. Language model output is open-ended and non-deterministic, so equality is almost never the right assertion. Two different answers can both be correct, and the same input can produce different text on Tuesday.
That doesn't mean you can't test it. It means the assertion moves from is it this exact string to does it satisfy this property. Most of the craft is picking properties precise enough to be checkable and important enough to be worth checking.
Offline and online evaluation
These answer different questions and neither replaces the other.
Offline
before you ship
Run a fixed set of examples against a change and compare. Answers is this version better than the last one?
Weakness: it only contains failures someone thought to write down.
Online
after you ship
Score real production traffic as it happens. Answers is this working for actual users right now?
Weakness: it tells you after the fact, unless it feeds back into a gate.
The loop that matters connects them: online evaluation finds a failure you never imagined, and that failure becomes an offline test so it can never ship again. Most teams run both halves and never join them, which is why the same bug ships twice.
Three ways to score, cheapest first
1. Deterministic checks
Is the JSON valid? Does it match the schema? Are required fields present, forbidden strings absent, latency and cost inside budget? These are free, instant, perfectly reliable, and catch a genuinely large share of real failures. Exhaust them before reaching for a model.
2. LLM-as-a-judge
A model scores the output against a rubric you write. This is the only practical way to measure groundedness, faithfulness to retrieved context, tone, or whether a task was completed. It costs money and tokens per evaluation, and it is itself a model that can be wrong — which is why the calibration step below is not optional. Full treatment, including the five biases that break judges and the rubric patterns that survive them, in our LLM-as-a-judge guide.
3. Human review
The ground truth everything else approximates. Too slow to run on everything, and that's fine: its highest-value use isn't grading output, it's grading your judges.
A note on classic reference metrics — BLEU, ROUGE, exact match. They compare against a reference answer and punish valid rewordings. For agent output they correlate poorly with whether the answer was useful. They still have a place in narrow summarisation and translation work.
The level you evaluate at decides what you can catch
This is the part most tooling gets wrong, and it's the difference between catching your real failures and catching the easy ones.
Did the retriever return relevant documents? Was the tool called with valid arguments?
Is the response grounded in what was retrieved? Does it follow the format contract?
Was the user's problem actually resolved across six turns — or did the agent claim success after a tool error?
An agent that calls a refund API, gets a timeout, and then tells the customer their refund was processed has no bad span. The API call was correct. The response was fluent. The failure exists only in the sequence — and a judge scoped to one observation cannot see it, no matter how good the rubric is. Check whether your tooling can target a conversation before you trust a green dashboard.
Then evaluate the evaluator
A judge you haven't checked is a confident random number generator. The fix is unglamorous: label a sample of runs by hand, compare against what the judge said, and measure agreement. Two error types matter and they cost differently — false passes are bugs reaching users, and false fails are an alert everyone learns to ignore.
Do this before a judge is allowed to block a deploy. A gate built on an uncalibrated judge is worse than no gate: it converts an unknown risk into false confidence.
Tools for running LLM evaluation
We build the first one, so weigh it accordingly — the specifics are checkable, which is the point.
Evaluators are columns on the trace table, judged at conversation, run or span level, running online as traces land and on demand. Failing production runs freeze into hermetic regression cases that replay in CI with no model spend. Judge calibration scores your evaluators against human labels, so you can tell whether the judge is right before you let it gate a release.
Cost $0 self-hosted with every feature and no paywalled internals. Free hosted tier at 20k traces/month; Team at $49/month. LLM judges run on your own model key — no markup on inference.
The most widely adopted option, and the biggest community in the category. Tracing, prompt management with versioning, evaluators, datasets and experiments in one mature product.
Cost Open-source core is free to self-host; some enterprise features are commercially licensed. Hosted plans available.
OpenInference-native tracing with strong notebook-driven analysis. A good fit if your evaluation work happens in Jupyter rather than in a dashboard.
Cost Phoenix is free and open source; the deeper platform features are in the commercial Arize product.
Libraries rather than platforms — you get evaluation metrics you can call from pytest, with no backend to run. Ragas is RAG-focused; DeepEval is broader.
Cost Free. You supply the storage, the dashboards and the production wiring yourself.
Both are polished commercial platforms with strong experiment workflows. LangSmith is the natural choice if your stack is LangChain end to end.
Cost Paid, hosted. Self-hosting is limited or unavailable depending on tier.
If you want the fuller head-to-head, we wrote one: Langfuse alternatives, compared honestly.
A starting checklist
- 1Trace everything first. You cannot evaluate what you didn't record, and instrumentation is the only step with no shortcut.
- 2Add deterministic checks: schema, required fields, latency, cost. Free, and they catch more than you'd expect.
- 3Write two or three judges for the properties you actually care about. Not twelve.
- 4Run them online, on production traffic, not only on a static set.
- 5Calibrate against human labels before any judge is allowed to block anything.
- 6Turn every real production failure into an offline test, so it cannot ship twice.
- 7Only then wire the gate into CI.
Run all seven steps for $0
Tracely is MIT-licensed with no paywalled internals — self-host the whole product free, or use the hosted free tier. Judges run on your own model key, so we never take a cut of inference.