Error Span
Build MQM-style error annotation interfaces in Potato for translation quality evaluation and typed error span marking with weighted severity scoring.
The error span annotation schema provides an MQM-style (Multidimensional Quality Metrics) interface for marking errors in text with typed categories and severity levels. Use it for translation quality evaluation, text editing review, content quality assessment, and any task requiring fine-grained error annotation.
Error span in Potato
Overview
The error span schema provides:
- Typed error categories with optional subtypes for detailed classification
- Severity levels with configurable point deductions
- Running quality score that decreases as errors are marked
- Color-coded spans that visually distinguish error types and severities
Annotators select text spans, assign an error type and severity, and the system automatically computes a quality score.
Quick Start
error_types is required and Potato refuses to load a scheme whose list is empty, since the categories are what the annotator assigns. Keeping the top level short makes the choice fast:
annotation_schemes:
- annotation_type: error_span
name: translation_errors
description: Mark all errors in the translation below.
error_types:
- name: Accuracy
- name: Fluency
- name: Terminology
show_score: true
max_score: 100Configuration Options
| Field | Type | Default | Description |
|---|---|---|---|
annotation_type | string | Required | Must be "error_span" |
name | string | Required | Unique identifier for this schema |
description | string | Required | Instructions displayed to annotators |
error_types | array | Required | List of error type objects, each with name and optional subtypes array |
severities | array | [{name: "Minor", weight: -1}, {name: "Major", weight: -5}, {name: "Critical", weight: -10}] | List of severity levels with name and weight (point deduction) |
show_score | boolean | true | Display a running quality score |
max_score | integer | 100 | Starting quality score before deductions |
source_field | string | "" | Display field this schema annotates, when the item shows more than one text |
Examples
Translation Quality (MQM)
The default weights follow MQM practice, where a critical error costs ten times a minor one. Subtypes keep the first choice to four categories while still recording which kind of accuracy error occurred:
annotation_schemes:
- annotation_type: error_span
name: mqm_errors
description: >
Mark all errors in the machine translation.
Select the error span, choose a category and severity.
error_types:
- name: Accuracy
subtypes:
- Mistranslation
- Addition
- Omission
- Untranslated
- name: Fluency
subtypes:
- Grammar
- Spelling
- Punctuation
- Register
- name: Terminology
subtypes:
- Inconsistent
- Wrong Term
- name: Style
severities:
- name: Minor
weight: -1
- name: Major
weight: -5
- name: Critical
weight: -10
show_score: true
max_score: 100Content Editing Review
Editing review asks which changes are obligatory, so two severities carry the distinction and a 5:1 weight ratio separates them. show_score: false hides the running total, which suits a pass where the marked errors are the deliverable and a score would only anchor the reviewer:
annotation_schemes:
- annotation_type: error_span
name: editing_errors
description: Mark all issues that need editing in this article.
error_types:
- name: Factual Error
- name: Grammar
subtypes:
- Subject-Verb Agreement
- Tense
- Pronoun Reference
- name: Style
subtypes:
- Wordiness
- Passive Voice
- Jargon
- name: Formatting
severities:
- name: Suggestion
weight: -1
- name: Required Fix
weight: -5
show_score: falseCode Review Annotation
Three severities spread over a 1, 3, 10 range let a reviewer separate a preference from a defect from a blocker, and the gap to 10 keeps a single blocker visible in the score:
annotation_schemes:
- annotation_type: error_span
name: code_errors
description: Mark issues in this code snippet.
error_types:
- name: Bug
subtypes:
- Logic Error
- Off-by-One
- Null Reference
- name: Style
subtypes:
- Naming
- Formatting
- name: Security
subtypes:
- Injection
- Exposure
- name: Performance
severities:
- name: Nitpick
weight: -1
- name: Warning
weight: -3
- name: Blocker
weight: -10
max_score: 100
show_score: trueOutput Format
{
"translation_errors": {
"labels": {
"errors": [
{
"start": 12,
"end": 25,
"text": "incorrectly translated",
"type": "Accuracy",
"subtype": "Mistranslation",
"severity": "Major"
},
{
"start": 45,
"end": 52,
"text": "the the",
"type": "Fluency",
"subtype": "Grammar",
"severity": "Minor"
}
],
"score": 94
}
}
}An error carries the category in type, and subtype is an empty string when the chosen category has no subtypes. Potato subtracts the absolute weight of every marked error from max_score and floors the result at 0, so the two errors above score 100 - 5 - 1 = 94.
Best Practices
- Define clear error type boundaries - annotators should not struggle to decide between two error types; provide examples in the description
- Use subtypes for granularity - top-level types keep the interface simple while subtypes allow detailed analysis when needed
- Calibrate severity weights carefully - the weight ratios should reflect actual impact; a critical error should be meaningfully more costly than a minor one
- Set max_score relative to text length - for short texts, a lower max_score prevents single errors from having outsized impact
- Provide annotation guidelines - MQM-style annotation benefits greatly from detailed guidelines with examples of each error type and severity
Further Reading
- Span Annotation - General span labeling
- Span Linking - Linking related spans
- Quality Control - Attention checks and gold standards
For implementation details, see the source documentation.