Inter-annotator agreement calculator
Paste a CSV of labels and get Krippendorff's alpha, Cohen's kappa (unweighted and weighted), Fleiss' kappa, and raw agreement, with a bootstrap 95% confidence interval on alpha. Agreement is also broken down by label and by annotator pair, which points to the label or the annotator behind low agreement. The computation runs in your browser and your data stays on your machine.
Computed locally in your browser. Your data is not uploaded anywhere.
Which coefficient should you report?
Krippendorff's alpha is the most general choice. It handles any number of annotators, missing ratings, and nominal, ordinal, or interval data. Report it when annotators did not all rate every item, which is the normal situation in crowdsourced studies. Cohen's kappa (Cohen, 1960) applies to exactly two annotators, and Fleiss' kappa (Fleiss, 1971) extends chance correction to groups where every item receives the same number of ratings. When your data violates those assumptions the calculator leaves the cell blank instead of producing a misleading number.
Interpretation conventions vary by field, but Krippendorff's own guidance is common: α ≥ 0.8 supports firm conclusions, 0.667 to 0.8 supports tentative ones, and anything lower means the labels are not reliable enough to analyze. Raw percent agreement is reported alongside these coefficients, never instead of them, because it ignores agreement that happens by chance.
For ordered labels such as a 1 to 5 rating, the Cohen's kappa card also reports weighted kappa (Cohen, 1968), which gives partial credit when two annotators pick neighboring values. Linear weights grow with the distance between the two values and quadratic weights with its square, so quadratic-weighted kappa forgives near misses more. The weights use rank positions, the convention of scikit-learn's cohen_kappa_score, which means a scale with gaps such as 1, 2, 5 is treated as evenly spaced. Set the data level for alpha to ordinal on the same data so that both coefficients give credit for near misses.
Agreement by label and by annotator pair
A single coefficient says how much annotators disagree but not where. For each label, the by-label table reports specific agreement and alpha for that label against all the others. Specific agreement is the share of within-item rating pairs involving the label in which the other rating is the same label. A label whose specific agreement sits far below the rest points to a definition the guidelines leave unclear. With three or more annotators, the pair table lists Cohen's kappa for every pair that rated the same items (the 25 weakest when there are more), lowest first, and Light's kappa (Light, 1971), the mean of those pairwise values, summarizes them. When one annotator appears in every weak pair, retrain or replace that annotator before revising the guidelines.
The implementation follows Krippendorff (2011), "Computing Krippendorff's Alpha-Reliability," and matches the simpledorff and krippendorff Python packages to six decimal places on shared test data. The weighted and pairwise kappas match scikit-learn's cohen_kappa_score to twelve decimal places on 40 randomly generated datasets. The confidence interval is a percentile bootstrap over items with 1000 resamples.
Range of alpha and its relation to Fleiss' kappa
Can Krippendorff's alpha be less than −1?
No, for nominal, ordinal, and interval data alpha cannot reach −1. Krippendorff (2011) gives alpha's range for reliability purposes as 1 ≥ α ≥ 0 and attributes values below zero to sampling error and systematic disagreement, without stating a floor. The floor follows from the formula. With two annotators per item, alpha is at least −1 + 2/n, where n is the number of ratings on items that have at least two, so four items rated by two annotators can go no lower than −0.75. Reaching that minimum requires the annotators to disagree the same way on every item, as when one always answers yes and the other always answers no. With m annotators on every item the floor rises to 1 − m(n − 1)/((m − 1)n), about −0.5 for three annotators. If software reports alpha below −1 on these data levels, check the input layout and the data level setting before reading anything into the number.
What is the difference between Fleiss' kappa and Krippendorff's alpha?
Fleiss' kappa requires every item to receive the same number of ratings and treats labels as unordered, while Krippendorff's alpha accepts missing ratings and supports ordinal and interval scales. On complete nominal data the two are linked exactly by α = κ + (1 − κ)/n, where κ is Fleiss' kappa and n is the total number of ratings. Alpha's small-sample correction accounts for the difference, which shrinks as the study grows, and with 300 ratings and κ = 0.6 alpha is 0.601. Report alpha when any annotator skipped items or the labels are ordered, and add Fleiss' kappa when prior work in your field reports it.
Related reading
The guide on measuring inter-annotator agreement covers study design choices, and agreement for span annotations explains why token-level metrics need different treatment. Potato computes kappa and alpha live in its admin dashboard during a study, so you can catch low agreement before the budget is spent.