Skip to content
Showcase/WMT15 Relative Ranking of Machine Translations
intermediatecomparison

WMT15 Relative Ranking of Machine Translations

The classic WMT relative-ranking (RR) protocol for machine translation evaluation, from 'Findings of the 2015 Workshop on Statistical Machine Translation' by Bojar et al. (WMT 2015). The annotator sees a source sentence, a human reference translation, and the outputs of up to five anonymized MT systems in random order, and ranks the outputs from best to worst with ties allowed; the joint rankings are reduced to pairwise judgments and aggregated with TrueSkill into the official system ranking.

About this dataset

From 2007 through 2016, the official ranking of machine translation systems at the WMT shared tasks was produced by human relative-ranking judgments. The 2015 findings paper by Bojar and colleagues documents the largest such campaign to that date: researcher judges ranked anonymized system outputs collected across ten translation tasks between English and each of Czech, French, German, Finnish, and Russian.

In each ranking task, an annotator is presented with a source segment, a human reference translation, and the outputs of up to five anonymized candidate systems, randomly selected and displayed in random order. The instruction shown at the top of each HIT reads: 'You are shown a source sentence followed by several candidate translations. Your task is to rank the translations from best to worst (ties are allowed).'

Annotators rank the outputs from 1 (best) to 5 (worst), with ties permitted. Each joint ranking is expanded into up to ten pairwise judgments (A > B, A = F, ...), and the pairwise judgments are aggregated with TrueSkill into a partial ordering that clusters systems by quality. In 2015, redundancy cleanup of near-identical outputs let the campaign nearly double its yield, collecting 542,732 annotations.

This Potato config reproduces a single WMT ranking task per instance using the drag-and-drop ranking scheme with ties enabled: the source and reference are displayed above five candidate translations drawn from the instance data. The sample items are self-authored illustrations of typical MT quality variation, not sentences from the WMT test sets.

Use this template for any evaluation where a small set of system outputs for the same input should be totally ordered — MT hypotheses, summaries, or LLM responses — when full rankings are preferred over pairwise or best-worst designs.

Campaign
WMT 2015 (Lisbon)
Protocol
Relative ranking (RR), ties allowed
Annotations collected
542,732
Candidates per task
Up to 5 anonymized system outputs
Translation tasks
10 (English ↔ Czech, French, German, Finnish, Russian)
Aggregation
TrueSkill over expanded pairwise judgments
Submit

Configuration Fileconfig.yaml

This Potato config reproduces the annotation task. Save it as config.yaml and run potato start config.yaml to try it.

yaml
# WMT15 Relative Ranking (RR) — Machine Translation Human Evaluation
# Based on: Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Barry Haddow,
#   Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz,
#   Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, and Marco Turchi (2015),
#   "Findings of the 2015 Workshop on Statistical Machine Translation."
#   Proceedings of the Tenth Workshop on Statistical Machine Translation (WMT 2015),
#   Lisbon, Portugal, pp. 1-46.
#   Paper: https://aclanthology.org/W15-3001/
#   Human judgment data: http://www.statmt.org/wmt15/results.html
#
# Task: the classic WMT relative-ranking (RR) protocol for machine translation
# evaluation. The annotator is shown a source segment, a human reference
# translation, and the outputs of up to five anonymized candidate systems in
# random order, and ranks the outputs from best (rank 1) to worst (rank 5), with
# ties permitted. Rankings are reduced to pairwise judgments and aggregated with
# TrueSkill to produce the official system ranking. Deliberate simplification: in
# WMT15 each HIT bundled three ranking tasks and near-identical outputs were
# merged into "multi-system translations"; here each Potato instance is a single
# ranking task over exactly five distinct candidate translations.
#
# Annotation instructions reproduced verbatim from Section 3.2 ("Data collection")
# of the paper, which quotes the instruction text provided at the top of each
# Appraise HIT; the practical notes below the quote are adapted from the
# surrounding protocol description in the same section.

annotation_task_name: "WMT15 Relative Ranking of Machine Translations"
task_dir: "."

data_files:
  - sample-data.json

item_properties:
  id_key: "id"
  text_key: "source"

output_annotation_dir: "annotation_output/"
output_annotation_format: "json"

port: 8000
server_name: localhost

annotation_instructions: |
  > You are shown a source sentence followed by several candidate translations.
  > Your task is to rank the translations from best to worst (ties are allowed).

  (Instruction reproduced verbatim from Section 3.2 of Bojar et al., 2015 — the
  instructions provided at the top of each HIT in the WMT15 evaluation campaign.)

  Practical notes, following the WMT15 protocol: drag the candidate translations
  so that the best translation is at the top and the worst is at the bottom; a
  lower rank is better. Use the tie control to place equally good translations at
  the same rank. The human reference translation is shown to help you judge
  adequacy. Skip a ranking task only as a last resort, e.g., if the translation
  candidates shown on screen are clearly misformatted or contain data issues
  (wrong language or similar problems).

annotation_schemes:
  - annotation_type: ranking
    name: translation_ranking
    description: "Rank the candidate translations from best (top) to worst (bottom). Ties are allowed."
    items_key: translations
    allow_ties: true
    label_requirement:
      required: true

html_layout: |
  <div style="padding: 15px; max-width: 860px; margin: auto;">
    <div style="background: #f1f5f9; border: 1px solid #cbd5e1; border-radius: 8px; padding: 6px 12px; margin-bottom: 8px; font-size: 13px; color: #475569;">
      <strong>Language pair:</strong> {{language_pair}}
    </div>
    <div style="background: #eff6ff; border: 1px solid #bfdbfe; border-radius: 8px; padding: 16px; margin-bottom: 8px;">
      <strong style="color: #1e40af;">Source sentence:</strong>
      <p style="font-size: 17px; line-height: 1.6; margin: 8px 0 0 0;">{{source}}</p>
    </div>
    <div style="background: #f0fdf4; border: 1px solid #bbf7d0; border-radius: 8px; padding: 16px; margin-bottom: 8px;">
      <strong style="color: #166534;">Human reference translation:</strong>
      <p style="font-size: 15px; line-height: 1.6; margin: 8px 0 0 0;">{{reference}}</p>
    </div>
    <div style="color: #6b7280; font-size: 13px;">
      The candidate system translations appear below in random order. Drag them to
      rank from best (top) to worst (bottom); ties are allowed.
    </div>
  </div>

allow_all_users: true
instances_per_annotator: 30
annotation_per_instance: 3
allow_skip: true

Sample Datasample-data.json

json
[
  {
    "id": "wmt15rr_001",
    "language_pair": "German → English",
    "source": "Die Bundesregierung kündigte am Dienstag ein neues Förderprogramm für kleine und mittlere Unternehmen an.",
    "reference": "The federal government announced a new support programme for small and medium-sized enterprises on Tuesday.",
    "translations": [
      "The federal government announced a new funding programme for small and medium-sized enterprises on Tuesday.",
      "On Tuesday the federal government has announced a new promotion program for small and middle enterprises.",
      "The federal government announced a new funding programme for enterprises.",
      "The government federal announced Tuesday new program of support for small and middle-standing companies.",
      "The Federal Government announced on Tuesday a new subsidy programme for small and medium companies."
    ]
  },
  {
    "id": "wmt15rr_002",
    "language_pair": "German → English",
    "source": "Wegen des anhaltenden Regens mussten mehrere Straßen in der Innenstadt gesperrt werden.",
    "reference": "Several roads in the city centre had to be closed because of the persistent rain.",
    "translations": [
      "Because of the continuing rain, several streets in the city centre had to be closed.",
      "Due to the persistent rain several roads in the inner city must be blocked.",
      "Because of the rain lasting, several streets in the downtown had to become blocked.",
      "Several streets in the city center were closed due to ongoing rain.",
      "Because of the persistent rain must several roads in the inner city be locked."
    ]
  }
]

// ... and 8 more items

Get This Design

View on GitHub

Clone or download from the repository

Quick start:

git clone https://github.com/davidjurgens/potato-showcase.git
cd potato-showcase/evaluation/wmt15-relative-ranking
potato start config.yaml

Dataset & paper

Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Barry Haddow, Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, and Marco Turchi. 2015. Findings of the 2015 Workshop on Statistical Machine Translation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 1-46, Lisbon, Portugal.

Citation (BibTeX)

bibtex
@inproceedings{bojar-etal-2015-findings,
    title = "Findings of the 2015 Workshop on Statistical Machine Translation",
    author = "Bojar, Ond{\v{r}}ej  and
      Chatterjee, Rajen  and
      Federmann, Christian  and
      Haddow, Barry  and
      Huck, Matthias  and
      Hokamp, Chris  and
      Koehn, Philipp  and
      Logacheva, Varvara  and
      Monz, Christof  and
      Negri, Matteo  and
      Post, Matt  and
      Scarton, Carolina  and
      Specia, Lucia  and
      Turchi, Marco",
    booktitle = "Proceedings of the Tenth Workshop on Statistical Machine Translation",
    month = sep,
    year = "2015",
    address = "Lisbon, Portugal",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/W15-3001/",
    doi = "10.18653/v1/W15-3001",
    pages = "1--46"
}

Details

Annotation Types

ranking

Domain

Machine TranslationNLP

Use Cases

MT System EvaluationHuman Evaluation CampaignsSystem Comparison

Tags

machine-translationrelative-rankingwmttrueskillhuman-evaluationsystem-ranking

Found an issue or want to improve this design?

Open an Issue