WMT15 Relative Ranking of Machine Translations
The classic WMT relative-ranking (RR) protocol for machine translation evaluation, from 'Findings of the 2015 Workshop on Statistical Machine Translation' by Bojar et al. (WMT 2015). The annotator sees a source sentence, a human reference translation, and the outputs of up to five anonymized MT systems in random order, and ranks the outputs from best to worst with ties allowed; the joint rankings are reduced to pairwise judgments and aggregated with TrueSkill into the official system ranking.
About this dataset
From 2007 through 2016, the official ranking of machine translation systems at the WMT shared tasks was produced by human relative-ranking judgments. The 2015 findings paper by Bojar and colleagues documents the largest such campaign to that date: researcher judges ranked anonymized system outputs collected across ten translation tasks between English and each of Czech, French, German, Finnish, and Russian.
In each ranking task, an annotator is presented with a source segment, a human reference translation, and the outputs of up to five anonymized candidate systems, randomly selected and displayed in random order. The instruction shown at the top of each HIT reads: 'You are shown a source sentence followed by several candidate translations. Your task is to rank the translations from best to worst (ties are allowed).'
Annotators rank the outputs from 1 (best) to 5 (worst), with ties permitted. Each joint ranking is expanded into up to ten pairwise judgments (A > B, A = F, ...), and the pairwise judgments are aggregated with TrueSkill into a partial ordering that clusters systems by quality. In 2015, redundancy cleanup of near-identical outputs let the campaign nearly double its yield, collecting 542,732 annotations.
This Potato config reproduces a single WMT ranking task per instance using the drag-and-drop ranking scheme with ties enabled: the source and reference are displayed above five candidate translations drawn from the instance data. The sample items are self-authored illustrations of typical MT quality variation, not sentences from the WMT test sets.
Use this template for any evaluation where a small set of system outputs for the same input should be totally ordered — MT hypotheses, summaries, or LLM responses — when full rankings are preferred over pairwise or best-worst designs.
- Campaign
- WMT 2015 (Lisbon)
- Protocol
- Relative ranking (RR), ties allowed
- Annotations collected
- 542,732
- Candidates per task
- Up to 5 anonymized system outputs
- Translation tasks
- 10 (English ↔ Czech, French, German, Finnish, Russian)
- Aggregation
- TrueSkill over expanded pairwise judgments
Configuration Fileconfig.yaml
This Potato config reproduces the annotation task. Save it as config.yaml and run potato start config.yaml to try it.
# WMT15 Relative Ranking (RR) — Machine Translation Human Evaluation
# Based on: Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Barry Haddow,
# Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz,
# Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, and Marco Turchi (2015),
# "Findings of the 2015 Workshop on Statistical Machine Translation."
# Proceedings of the Tenth Workshop on Statistical Machine Translation (WMT 2015),
# Lisbon, Portugal, pp. 1-46.
# Paper: https://aclanthology.org/W15-3001/
# Human judgment data: http://www.statmt.org/wmt15/results.html
#
# Task: the classic WMT relative-ranking (RR) protocol for machine translation
# evaluation. The annotator is shown a source segment, a human reference
# translation, and the outputs of up to five anonymized candidate systems in
# random order, and ranks the outputs from best (rank 1) to worst (rank 5), with
# ties permitted. Rankings are reduced to pairwise judgments and aggregated with
# TrueSkill to produce the official system ranking. Deliberate simplification: in
# WMT15 each HIT bundled three ranking tasks and near-identical outputs were
# merged into "multi-system translations"; here each Potato instance is a single
# ranking task over exactly five distinct candidate translations.
#
# Annotation instructions reproduced verbatim from Section 3.2 ("Data collection")
# of the paper, which quotes the instruction text provided at the top of each
# Appraise HIT; the practical notes below the quote are adapted from the
# surrounding protocol description in the same section.
annotation_task_name: "WMT15 Relative Ranking of Machine Translations"
task_dir: "."
data_files:
- sample-data.json
item_properties:
id_key: "id"
text_key: "source"
output_annotation_dir: "annotation_output/"
output_annotation_format: "json"
port: 8000
server_name: localhost
annotation_instructions: |
> You are shown a source sentence followed by several candidate translations.
> Your task is to rank the translations from best to worst (ties are allowed).
(Instruction reproduced verbatim from Section 3.2 of Bojar et al., 2015 — the
instructions provided at the top of each HIT in the WMT15 evaluation campaign.)
Practical notes, following the WMT15 protocol: drag the candidate translations
so that the best translation is at the top and the worst is at the bottom; a
lower rank is better. Use the tie control to place equally good translations at
the same rank. The human reference translation is shown to help you judge
adequacy. Skip a ranking task only as a last resort, e.g., if the translation
candidates shown on screen are clearly misformatted or contain data issues
(wrong language or similar problems).
annotation_schemes:
- annotation_type: ranking
name: translation_ranking
description: "Rank the candidate translations from best (top) to worst (bottom). Ties are allowed."
items_key: translations
allow_ties: true
label_requirement:
required: true
html_layout: |
<div style="padding: 15px; max-width: 860px; margin: auto;">
<div style="background: #f1f5f9; border: 1px solid #cbd5e1; border-radius: 8px; padding: 6px 12px; margin-bottom: 8px; font-size: 13px; color: #475569;">
<strong>Language pair:</strong> {{language_pair}}
</div>
<div style="background: #eff6ff; border: 1px solid #bfdbfe; border-radius: 8px; padding: 16px; margin-bottom: 8px;">
<strong style="color: #1e40af;">Source sentence:</strong>
<p style="font-size: 17px; line-height: 1.6; margin: 8px 0 0 0;">{{source}}</p>
</div>
<div style="background: #f0fdf4; border: 1px solid #bbf7d0; border-radius: 8px; padding: 16px; margin-bottom: 8px;">
<strong style="color: #166534;">Human reference translation:</strong>
<p style="font-size: 15px; line-height: 1.6; margin: 8px 0 0 0;">{{reference}}</p>
</div>
<div style="color: #6b7280; font-size: 13px;">
The candidate system translations appear below in random order. Drag them to
rank from best (top) to worst (bottom); ties are allowed.
</div>
</div>
allow_all_users: true
instances_per_annotator: 30
annotation_per_instance: 3
allow_skip: true
Sample Datasample-data.json
[
{
"id": "wmt15rr_001",
"language_pair": "German → English",
"source": "Die Bundesregierung kündigte am Dienstag ein neues Förderprogramm für kleine und mittlere Unternehmen an.",
"reference": "The federal government announced a new support programme for small and medium-sized enterprises on Tuesday.",
"translations": [
"The federal government announced a new funding programme for small and medium-sized enterprises on Tuesday.",
"On Tuesday the federal government has announced a new promotion program for small and middle enterprises.",
"The federal government announced a new funding programme for enterprises.",
"The government federal announced Tuesday new program of support for small and middle-standing companies.",
"The Federal Government announced on Tuesday a new subsidy programme for small and medium companies."
]
},
{
"id": "wmt15rr_002",
"language_pair": "German → English",
"source": "Wegen des anhaltenden Regens mussten mehrere Straßen in der Innenstadt gesperrt werden.",
"reference": "Several roads in the city centre had to be closed because of the persistent rain.",
"translations": [
"Because of the continuing rain, several streets in the city centre had to be closed.",
"Due to the persistent rain several roads in the inner city must be blocked.",
"Because of the rain lasting, several streets in the downtown had to become blocked.",
"Several streets in the city center were closed due to ongoing rain.",
"Because of the persistent rain must several roads in the inner city be locked."
]
}
]
// ... and 8 more itemsGet This Design
Clone or download from the repository
Quick start:
git clone https://github.com/davidjurgens/potato-showcase.git cd potato-showcase/evaluation/wmt15-relative-ranking potato start config.yaml
Dataset & paper
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Barry Haddow, Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, and Marco Turchi. 2015. Findings of the 2015 Workshop on Statistical Machine Translation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 1-46, Lisbon, Portugal.
Citation (BibTeX)
@inproceedings{bojar-etal-2015-findings,
title = "Findings of the 2015 Workshop on Statistical Machine Translation",
author = "Bojar, Ond{\v{r}}ej and
Chatterjee, Rajen and
Federmann, Christian and
Haddow, Barry and
Huck, Matthias and
Hokamp, Chris and
Koehn, Philipp and
Logacheva, Varvara and
Monz, Christof and
Negri, Matteo and
Post, Matt and
Scarton, Carolina and
Specia, Lucia and
Turchi, Marco",
booktitle = "Proceedings of the Tenth Workshop on Statistical Machine Translation",
month = sep,
year = "2015",
address = "Lisbon, Portugal",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/W15-3001/",
doi = "10.18653/v1/W15-3001",
pages = "1--46"
}Details
Annotation Types
Domain
Use Cases
Tags
Found an issue or want to improve this design?
Open an IssueRelated Designs
EA-MT - Entity-Aware Machine Translation
Entity-aware machine translation evaluation requiring annotators to identify entity spans, classify translation errors, and provide corrected translations. Based on SemEval-2025 Task 2.
ESA: Error Span Annotation for Machine Translation
Error span annotation for machine translation output. Annotators identify error spans in translations, classify error types (accuracy, fluency, terminology, style), and rate severity.
FLORES - Machine Translation Quality Estimation
Machine translation quality assessment using the FLORES-101 benchmark (Goyal et al., TACL 2022). Annotators rate translation quality on a Likert scale, identify error categories, and provide detailed error notes.