Dawid-Skene consensus label calculator
Aggregate noisy labels from multiple annotators into consensus labels with the Dawid-Skene model or MACE, and see the reliability each model estimates for every annotator. The computation runs in your browser and your data stays on your machine.
Computed locally in your browser. Your data is not uploaded anywhere.
Why not majority vote?
Majority voting weights every annotator equally. In practice some annotators are careful, some misunderstand one specific label, and a few click through at random. The Dawid-Skene model (Dawid & Skene, 1979) estimates a confusion matrix for each annotator and the true label of each item jointly, using expectation-maximization. An annotator who is usually right gets more say; one who systematically confuses two labels gets corrected rather than discarded. On items where the model and the majority disagree, the model is using reliability evidence the vote count cannot see.
The estimated accuracy column doubles as quality control. It flags unreliable annotators without requiring gold-standard questions. It only works when annotators overlap on shared items, so plan at least three ratings per item if you intend to use it.
Dawid-Skene and MACE compared
MACE (Hovy et al., 2013) models each annotator more simply than Dawid-Skene. On each item the annotator either knows the true label and reports it, or guesses from a personal distribution over labels, and the model estimates the probability of knowing, which it calls competence. Dawid-Skene fits a full confusion matrix per annotator instead, so it can learn that one annotator calls every neutral item negative and correct for that habit, at the cost of L × (L − 1) free parameters per annotator for L labels against MACE's L. When each annotator labels only a few dozen items, MACE's smaller model is less likely to overfit, because it estimates L parameters per annotator instead of L × (L − 1). Its competence score also separates spammers, since an annotator who always picks the same label gets a competence near zero.
The calculator's MACE runs EM with the true-label distribution fixed to uniform, 10 random restarts of 50 iterations each, and keeps the restart with the highest log-likelihood. On a simulated study of 300 items and five annotators whose true competence was 0.9, 0.85, 0.6, 0 and 0, MACE estimated 0.93, 0.83, 0.53, 0.00 and 0.00. MACE's consensus labels were correct on 93% of items and Dawid-Skene's on 92%, against 86% for majority vote. The summary line above the tables counts the items where Dawid-Skene and MACE disagree, and those items are the first ones to send to adjudication.
Related reading
The guide on aggregating crowd labels compares majority vote, Dawid-Skene, and newer aggregation models, and adjudication and disagreement covers what to do with the items the model is unsure about. Potato runs the same Dawid-Skene model when it aggregates annotations into an evaluation dataset. Potato can also run MACE during a study, recomputing competence as annotations arrive, so a likely spammer is flagged while the study is still running. Potato's MACE puts a Beta prior on spamming (alpha 0.5 by default), so its competence estimates can differ from the ones here.