Skip to content

Best-worst scaling score calculator

Turn best-worst scaling judgments into a score for every item. Paste one row per judgment with the items the annotator saw and the ones chosen as best and worst, and the calculator reports counting scores, a Bradley-Terry fit, and split-half reliability. The computation runs in your browser and your data stays on your machine.

Computed locally in your browser. Your data is not uploaded anywhere.

Counting scores and the Bradley-Terry fit

In best-worst scaling an annotator sees a small set of items, usually four, and picks the one with the most and the one with the least of some property, such as the most and least positive word. The counting score of an item is the share of its appearances in which it was chosen best minus the share in which it was chosen worst, so it runs from −1 to 1. Counting is the procedure Kiritchenko and Mohammad (2017) used to score 3,207 English terms for sentiment intensity, and it is the easiest score to check by hand.

The calculator fits a Bradley-Terry model to the pairs each judgment implies. The best item beat every other item in its tuple and every other item beat the worst, which gives five pairs per judgment of four items, and the Bradley-Terry model turns those pairs into a strength per item. Unlike counting, the fit takes into account which items an item beat and which it lost to. In one simulation of 24 items with known scores, 48 tuples, and three judgments per tuple, counting scores correlated with the true scores at Spearman ρ = 0.86 and Bradley-Terry strengths at 0.90.

Split-half reliability

Split-half reliability checks whether the ranking would come out the same from a different set of judgments. Following Kiritchenko and Mohammad (2017), the calculator splits the judgments of every tuple at random into two halves, scores each half by counting, and computes Spearman's ρ between the two sets of scores, averaged over 100 random splits. Tuples judged only once are dealt evenly between the halves. Their study generated 2N distinct 4-tuples for N terms, so that each term appeared in eight tuples, and had each tuple judged by ten annotators. That design reached a split-half reliability of 0.98, against 0.95 for a rating scale with the same number of annotations.

Each half holds only half of the judgments, so the number describes a study half the size of yours and understates the reliability of the full set. A low value with few judgments per tuple is a reason to collect more before trusting the ranking. In the same study, two judgments for each of 1.5N tuples matched the reliability a rating scale reached with ten ratings per term.

Related reading

The guide to pairwise comparison and best-worst scaling covers when to choose either design over a rating scale. To collect the judgments, Potato's best-worst scaling scheme generates the tuples and records best and worst, and python -m potato.bws_scoring --config config.yaml --method counting scores them from the study's output, with bradley_terry and plackett_luce as the other methods.