Guide · 2026
LLM-as-a-judge, and how to know it's right
Using a language model to score another model's output is the only practical way to measure the things that matter. It's also a model, which means it can be confidently wrong. Here's how it works, the biases that break it, and the step most teams skip.
What LLM-as-a-judge is
A judge is a language model given three things: the output to assess, criteria written in plain language, and a required response format. It returns a verdict — pass or fail, a score, a label — that gets attached to the run as a score you can filter, chart and alert on.
It exists because the interesting properties resist assertion. You can assert that JSON parses. You cannot assert that an answer is grounded in the retrieved documents, that the tone suits a frustrated customer, or that a six-step agent actually resolved the request. Human review can judge all three and doesn't scale past a sample. A model judge is the compromise: worse than a careful human, available on every single run.
The mechanics
- 1. Take the run's input, output, and any context it used.
- 2. Fill a rubric template with them.
- 3. Ask a model for a verdict in a fixed schema — reasoning first, verdict last.
- 4. Write the verdict back onto the run as a score.
Use it only for what you can't assert
A judge costs tokens, adds latency, and introduces a second source of error. Anything a deterministic check can catch should be caught by one: valid JSON, schema conformance, required fields present, forbidden strings absent, latency and cost inside budget. Those are free, instant and perfectly reliable.
Reach for a judge when the property is genuinely semantic — groundedness against retrieved context, faithfulness to a source document, tone, refusal appropriateness, task completion. Two or three good judges beat twelve mediocre ones, and every extra judge is another thing that can drift.
The five ways judges go wrong
These are well documented and they are not edge cases. Each has a cheap mitigation.
Position bias
In pairwise comparisons, judges favour whichever response came first.
Mitigation Run both orderings and keep the result only when they agree — or avoid pairwise entirely and score against an absolute rubric.
Verbosity bias
Longer, more confident-sounding answers score higher whether or not they are more correct.
Mitigation State length expectations in the rubric, and score correctness separately from completeness.
Self-preference
A judge rates output from its own model family more generously.
Mitigation Judge with a different family than the one under test. Never let a model grade its own homework.
Score clustering
On a 1–10 scale, almost everything lands on 7 or 8. The distribution carries little information.
Mitigation Use binary pass/fail, or decompose into several binary criteria and count passes.
Rubric sensitivity
Small rewordings of the prompt move verdicts more than real quality differences do.
Mitigation Version the rubric like code and re-check agreement whenever it changes.
Rubric patterns that hold up
- Reasoning before verdict. Make the judge state why, then decide. A verdict-first schema is a coin flip with a justification bolted on afterwards.
- One criterion per judge. “Is this helpful, accurate and well-formatted?” produces a verdict you can't act on. Three judges produce three you can.
- Define failure, not just success. Concrete examples of what a FAIL looks like move agreement more than any amount of describing PASS.
- Give it the context the run had. A groundedness judge without the retrieved documents is guessing, and will confidently tell you the answer was grounded.
- Pin the model and version the rubric. Both are inputs to your baseline. Changing either silently invalidates every historical comparison.
- Allow abstention. A judge that must answer will invent a verdict on a run it can't assess. Let it skip — and make sure your roll-up treats a skip as “not evaluated” rather than “passed”.
Judge the right thing
A rubric can be perfect and still miss the bug, because the judge was pointed at the wrong scope. An agent that calls a refund API, receives a timeout, and then tells the customer the refund went through has no bad span — the call was correct, the reply was fluent. The failure exists only in the sequence.
So check what your tooling can target before you trust a green dashboard. A judge scoped to a single observation cannot see a trajectory failure no matter how good the rubric is. Tracely judges at conversation, run or span level as a field on the evaluator; more on why the level decides what you catch in our guide to LLM evaluation, and how the tools differ in our Langfuse comparison.
The step almost everyone skips: calibrate the judge
Everything above improves a judge. None of it tells you whether the judge is right. For that there is exactly one method: label a sample of runs by hand, compare against what the judge said, and measure agreement.
False pass
The judge said PASS, the human said FAIL. A real bug reached a real user and your dashboard stayed green.
The expensive error. Optimise against this.
False fail
The judge said FAIL, the human said PASS. Over-flagging that trains the team to ignore the signal.
The insidious error. It kills the practice rather than the release.
Fifty labelled runs per evaluator is enough to see whether a judge is usable. Do it before a judge is allowed to block a deploy — a gate built on an uncalibrated judge is worse than no gate, because it converts an unknown risk into false confidence. Then re-check whenever you change the rubric or the model.
This is the part of the workflow we think is most underserved, so it's built into Tracely: label verdicts in the UI, get per-evaluator agreement with false-pass and false-fail broken out, and see an over-flagging judge before it gates anything.
A working checklist
- 1Exhaust deterministic checks first — a judge should never be asked whether JSON parses.
- 2Write two or three judges, one criterion each, reasoning before verdict.
- 3Use a judge model from a different family than the system under test, and pin it.
- 4Prefer binary verdicts over numeric scores.
- 5Point each judge at the right level — conversation, run or span.
- 6Label ~50 runs by hand and measure agreement before trusting it.
- 7Only then let a judge gate a release.
Judges, calibration and the gate — for $0
Tracely is MIT-licensed with no paywalled internals: self-host the whole product free, or use the hosted free tier. Judges run on your own model key, so we never take a cut of inference.