Skip to content

LLM judge agreement and calibration calculator

Check an LLM-as-a-judge against human labels. Paste the judge's labels together with those of people who labeled the same items without seeing the judge's labels, and the calculator compares the judge's agreement with the people to the people's agreement with each other. Two people give the human–human baseline, and the alternative annotator test needs three. When the judge reports a confidence, it also measures calibration. The computation runs in your browser and your data stays on your machine.

Computed locally in your browser. Your data is not uploaded anywhere.

Judge–human agreement against human–human agreement

Accuracy against one person's labels mixes the judge's errors with that person's, and on a subjective task no rater reaches 100%. The calculator therefore reports the mean Cohen's kappa between the judge and each human next to the mean kappa between pairs of humans, which is the agreement one more human annotator would be expected to reach on the same items. If two people agree at κ = 0.62, a judge at 0.60 with each of them agrees with them about as well as they agree with each other, and a judge at 0.35 is not, however good its raw accuracy looks. The gap card gives the difference with a 95% bootstrap interval over items, so when the interval spans zero, the data do not show that the judge is worse than a person.

The alternative annotator test (Calderon et al., 2025) asks whether the judge could replace one of the people. It leaves each human out in turn and, item by item, checks whether the judge or the left-out human matches the remaining humans more closely. For each human, a one-sided paired t-test checks whether the judge falls short of that human by less than a margin ε, measured as the share of items where each is at least as close to the remaining humans. The p-values are corrected across humans with Benjamini-Yekutieli. The judge passes when the test rejects for at least half of the humans, a winning rate ω ≥ 0.5. ε prices the cost of human labels, and the authors recommend 0.1 for crowd workers, 0.15 for skilled annotators, and 0.2 for experts. The test needs at least three humans, and the paper recommends 30 or more items per human. To compare two judges, use the advantage probability ρ, the judge's at-least-as-close share averaged over humans.

Calibration of the judge's confidence

A judge is calibrated when its confidence matches how often it is right, so that of the verdicts given at 80%, about 80% agree with the people. Expected calibration error (ECE) sorts verdicts into ten equal-width confidence bins and averages the gap between each bin's mean confidence and its accuracy, weighted by the number of verdicts in the bin. The Brier score is the mean squared difference between the confidence and whether the verdict was correct. Both use the human majority label as the reference and skip items where the humans tie. Bins whose accuracy falls more than ten points below their confidence are highlighted, because a filter that keeps only high-confidence verdicts keeps an overconfident judge's errors.

Confidence can be a probability between 0 and 1 or a percentage. If the judge's stated confidences cluster near the top of the scale, sample the judge several times per item and use the share of samples that agree with its most common label. Potato's judge calibration workflow computes confidence that way.

Ordered labels and the disagreement list

When every label is a number, such as a 1 to 5 quality rating, the pair table adds quadratic-weighted kappa, which counts a 4 against a 5 as a near miss rather than a full disagreement. The disagreement list sorts the items where the judge contradicts the human majority by the judge's confidence, so the confident errors come first. Those items are the ones to read before revising the judge's rubric, and the CSV download keeps the full list.

Related reading

The guide to LLM judge and human agreement covers how to collect the human sample, keep annotators blind to the judge, and test pairwise judges for position bias. To compare a judge with human labels inside a study, Potato's judge alignment shows the judge's verdict next to each human label with a running kappa when inline mode is on, and judge calibration reports ECE and Brier score against blind human labels. The inter-annotator agreement calculator reports Krippendorff's alpha when some raters skipped items.