Span, bounding box, and time segment agreement calculator
Measure inter-annotator agreement on bounding boxes in images, entity spans in text, and labeled segments on an audio or video timeline. Paste a CSV with one row per annotation, and the calculator matches each annotator's objects to the others' before reporting whether they found the same objects, labeled them the same way, and put them in the same place. The computation runs in your browser and your data stays on your machine.
Computed locally in your browser. Your data is not uploaded anywhere.
Chance-corrected agreement for bounding boxes
Mean IoU between annotators is a common agreement number for boxes, and it runs high on any dataset where objects are large and centered, even for an annotator who ignores the image. We simulated 50 images that each hold one large, roughly centered object. An annotator who drew the same fixed box on every image reached a mean matched IoU of 0.83 against a careful annotator, close to the 0.90 two careful annotators reached. The localization score σ separates the two cases, at −0.21 for the fixed box against 0.58 for the careful pair.
σ compares how far apart matched boxes are on the same image with how far apart boxes are on different images, σ = 1 − mean(within-item distance) / mean(between-item distance). At 1 the annotators agree perfectly, at 0 they agree no better than boxes from unrelated images, and below 0 they disagree systematically. The default distance is one minus generalized IoU, halved to the range 0 to 1, which still separates a near miss from a far one when boxes do not overlap, where plain IoU is zero for both. The Kolmogorov-Smirnov statistic compares the two distance distributions as a whole. Because σ compares means, a few images with very distant boxes can pull it down while KS stays high, so a high KS with a low σ points to a few images carrying most of the disagreement.
Location is only one of three ways annotators disagree. Detection α asks whether they found the same objects, scoring each cluster of overlapping boxes as present or absent per annotator, and classification α asks whether matched boxes got the same label. Both are nominal Krippendorff's alpha. High detection with low localization points to drawing guidelines that are too loose, and low detection with high localization points to an unclear definition of an object. Detection α can come out negative on small studies where nearly every object is found, because the few misses are all the variation alpha has to work with. The match threshold moves disagreement between detection and localization, since a pair of boxes at IoU 0.4 counts as two objects at the default 0.5 and as one loosely drawn object at 0.3.
Exact, partial, and character-level span agreement
Exact span F1 counts two spans as the same only when the start, end, and label all match. Partial F1 accepts a pair with the same label whose overlap covers at least half of one of the two spans, so "Obama" and "Barack Obama" still agree. Character α labels every character with its span label, or O (outside) for characters in no span, and computes nominal alpha over the characters of each item before averaging over items. Supply the item lengths for character α, since characters after the last marked span otherwise go uncounted.
The disagreement table splits every annotator pair's spans into five outcomes. Spans with the same boundaries and a different label point to a label definition problem, spans with the same label and different boundaries point to unclear rules about where an entity starts and ends, and missed spans point to disagreement about what counts as an entity at all. A single F1 adds these outcomes together, so it does not say which part of the guidelines to change. The guide to agreement for span annotations covers the choice of unit in more depth.
Time segments and the boundary tolerance sweep
Segments are matched by temporal IoU at the match threshold, 0.5 by default as for boxes, and the calculator reports the mean matched IoU, detection F1, and the detection and classification alphas over clusters of overlapping segments. It does not report σ for segments. A chance baseline built from segments of different clips assumes the clips are comparable, which fails when clip lengths differ.
Two annotators marking when speech starts will rarely pick the same frame, so the boundary table reports agreement at several tolerances instead of one. Boundary F1 at a tolerance is the share of segment start and end points that have a partner from the other annotator within that many seconds, matched one to one. High agreement at ±2 s with low agreement at ±0.1 s means the annotators agree an event happened but not exactly when, and the guidelines should define when it starts. Report the tolerance next to the number in a paper, because agreement at 0.25 s and at 2 s are different claims.
Comparison with Potato's agreement report
The box measures port the method Potato uses for its agreement over geometry, which builds on Braylan, Alonso, and Lease (2022). On 14 test cases generated by running the code of Potato 2.10.2, the calculator reproduces detection α, classification α, σ, the KS statistic, mean matched IoU, detection F1, the span F1 scores, and character α to floating-point precision, and it matches Potato's object matching on 400 random cases. When a dataset has more than 2,000 pairs of objects on different items, Potato samples 2,000 of them for the chance baseline while the calculator averages over all of them, so σ can differ slightly there. The span disagreement table, the boundary tolerance sweep, the segment alphas, and the bootstrap interval on σ are this calculator's additions.
Questions about spatial and temporal agreement
How do you calculate inter-annotator agreement for bounding boxes?
Match each annotator's boxes to the other annotator's boxes by IoU, then score detection, classification, and localization separately. Report the localization distance against a chance baseline built from boxes on different images, because raw IoU is high on any dataset where objects are large and centered. Paste the boxes into the calculator above with one row per box, using an empty row for an annotator who found nothing on an image, since a missing row reads as an image that annotator never saw.
What is temporal IoU?
Temporal IoU is the length of the overlap between two time segments divided by the length of their union, so segments from 0 to 4 s and from 1 to 5 s score 3/5 = 0.6. It is the one-dimensional version of the IoU used for bounding boxes. Annotation studies use it to match segments between annotators, with Potato counting a pair as the same event at 0.5 or above, and then report the mean IoU of the matched pairs alongside how many segments found a partner.
Related reading
The guide on measuring agreement on bounding boxes walks through study design for detection datasets, and the inter-annotator agreement calculator handles labels without a location. To collect boxes, spans, or segments, start from the Potato quick start. Potato's admin agreement report computes σ, KS, and both alphas for every image annotation scheme while the study runs.