Raw IoU Is Not Agreement
Mean IoU between annotators can stay near 0.95 even for one who never looks. Potato reports σ, which estimates the chance baseline from other images.
Raw IoU between annotators tracks how constrained a spatial task is, so on a corpus of large, centered objects it can stay near 0.95 even for an annotator who never looks. A chance-corrected coefficient measures agreement instead.
The computer-vision annotation platforms we have found that report inter-annotator agreement all report it as raw IoU (intersection over union) against a per-class threshold. CVAT's consensus engine does, and so does V7's consensus stage. Label Studio Enterprise's metric list is exact match, numeric difference, IoU, and span overlap. None of them applies a chance correction, and without one the number does not mean what it appears to mean.
Mean IoU on a corpus of centered objects
Take a detection corpus where each image holds one large, roughly centered object. Two annotators both draw a box around it, their boxes overlap heavily, mean IoU comes out around 0.95, and the dashboard says agreement is excellent. Now replace one annotator with a process that draws a box in the middle of every image without looking, and mean IoU stays around 0.95. The measure is reporting a property of the corpus (how constrained the task is) rather than a property of the annotators.
Percent agreement fails the same way on categorical labels when one class dominates, and that failure has been understood since Cohen introduced kappa in 1960, and the platforms above report spatial agreement without the equivalent correction.
Krippendorff's alpha over IoU distance
Krippendorff's α accepts any distance function, so the obvious move is α over 1 − IoU, and that was our original plan, but it fails for a structural reason. α computes 1 − D_o/D_e, where D_e is expected disagreement under random pairing. IoU distance is bounded in [0, 1], and two randomly paired boxes almost always have IoU 0. D_e therefore collapses to approximately 1, and the whole coefficient degenerates into 1 − mean observed distance, with no working chance correction inside it.
IoU is also flat exactly where annotators disagree, because two boxes that do not overlap score 0 whether they are touching or at opposite corners of the image. The measure has no gradient in the region you care about.
The objection is also empirical. Braylan, Alonso and Lease (WWW 2022) evaluated candidate distance functions by whether an agreement measure orders them the way prior work does, with distances that are finer-grained and more sensitive to label error ranked higher. For bounding boxes, α ranks plain L2 (0.687) above GIoU (0.507) and IoU (0.505), while their two distributional measures give the expected order of GIoU, then IoU, then L2.
The σ coefficient and its empirical chance baseline
We report σ instead, which keeps α's form but estimates the chance baseline empirically from annotations of other items. The name and the between-item baseline come from Braylan, Alonso and Lease, whose σ is the share of observed distances that fall significantly below the between-item distribution. Potato's σ is the simpler ratio of means:
σ = 1 − mean(within-item distance) / mean(between-item distance)Instead of assuming a distribution of random pairings, the denominator compares annotations of different items. That comparison gives an empirical answer to "how far apart would these annotators be if they were not looking at the same thing", which is what a chance baseline is supposed to be.
σ reads as follows:
| σ | Reading |
|---|---|
| 1.0 | Perfect agreement |
| 0.0 | No better than annotating unrelated images |
| below 0 | Systematically further apart on the same image than on different ones |
We do not clamp negative values. A negative σ is a real and diagnosable state, and it almost always means the guidelines are ambiguous rather than that the annotators are careless. Clamping would hide that signal.
Alongside σ we report a two-sample Kolmogorov–Smirnov (KS) statistic between the within-item and between-item distance distributions. σ compares two means and can be dragged by outliers, while KS compares whole distributions. When σ and KS disagree, a minority of items is carrying most of the disagreement, which tells you which items to open.
Detection, classification, and localization as separate numbers
Even correctly chance-corrected, a single spatial agreement figure cannot say which of three things went wrong:
- Detection. Did they find the same objects at all?
- Classification. Did they give matched objects the same label?
- Localization. Are matched objects in the same place?
An annotator who finds every object and mislabels them all is a completely different problem from one who labels correctly but boxes loosely, and the fixes differ. Potato therefore reports the three separately.
STAPLE for masks and rotated IoU for 3D cuboids
For pixel masks, "do they agree" and "whose boundary should the dataset record" are different questions. STAPLE answers the second by estimating a latent consensus boundary along with a sensitivity and specificity per annotator. The two are reported separately, because under-segmenting and over-segmenting have opposite fixes.
STAPLE also weights by demonstrated reliability rather than counting votes, which matters when careful annotators are outnumbered. Potato's STAPLE test sets up a 40 × 40 pixel image whose true foreground is a 20 × 20 square, traced exactly by two careful annotators, while each of three noisy annotators keeps about half the square's pixels and marks about 35% of the image at random. A three-of-five majority vote scores Dice 0.846 against that truth, and STAPLE scores 1.000.
For 3D cuboids, use exact rotated 3D IoU. An axis-aligned approximation is adequate for level automotive data and wrong for drone, handheld, and indoor scans, where it manufactures disagreement that is an artifact of the measure.
The image adjudication bug and the limits of σ
The bug came first. Potato's adjudication compared the keys of an annotation, and an image schema stores everything under one key called _data, so every image pair scored 1.0 agreement. Two annotators who agreed on nothing looked unanimous, and no image was ever routed for review.
Fixing the bug meant asking what the right number would have been. The chance baseline drawn from annotations of different items and the KS comparison both follow Braylan, Alonso and Lease. The measures have one stated limit. Temporal segments currently get the uncorrected measures only, because "a segment from a different clip" is not a meaningful chance comparison when clips differ in length.
Further reading
References
Works cited in this post, in order of first mention:
Jacob Cohen (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement. https://doi.org/10.1177/001316446002000104
Alexander Braylan, Omar Alonso & Matthew Lease (2022). Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation Tasks. Proceedings of the ACM Web Conference 2022. https://doi.org/10.1145/3485447.3512242
S. K. Warfield, K. H. Zou & W. M. Wells (2004). Simultaneous Truth and Performance Level Estimation (STAPLE): An Algorithm for the Validation of Image Segmentation. IEEE Transactions on Medical Imaging. https://doi.org/10.1109/TMI.2004.828354