Skip to content

Confidence Annotation

Add confidence ratings paired with other annotations in Potato using Likert scales or sliders to capture how certain an annotator was about a decision.

The confidence annotation schema lets annotators rate their confidence in another annotation they have made. It pairs a confidence scale (Likert or slider) with a target annotation schema, which records how certain the annotator was alongside the choice itself.

Potato confidence annotation pairing a sentiment label with a 5-point confidence scaleConfidence annotation in Potato

Overview

Confidence ratings give you a per-item measure of how hard the decision was, which is what identifies ambiguous items, supports weighting during aggregation, and flags candidates for adjudication. When configured, a confidence widget appears alongside the target annotation and asks the annotator how sure they are of their decision.

Quick Start

Point target_schema at the annotation the rating is about, so the pair can be read back together. A 5-point Likert scale is the default and the fastest for an annotator to answer:

yaml
annotation_schemes:
  - annotation_type: radio
    name: sentiment
    description: What is the sentiment of this text?
    labels: ["Positive", "Negative", "Neutral"]
 
  - annotation_type: confidence
    name: sentiment_confidence
    description: How confident are you in your sentiment label?
    target_schema: sentiment
    scale_type: likert
    scale_points: 5

Configuration Options

FieldTypeDefaultDescription
annotation_typestringRequiredMust be "confidence"
namestringRequiredUnique identifier for this schema
descriptionstringRequiredInstructions displayed to annotators
target_schemastringOptionalName of the annotation schema this confidence rating applies to
scale_typestring"likert"Type of scale: "likert" for discrete points or "slider" for continuous
scale_pointsinteger5Number of points on the Likert scale (ignored for slider)
labelsarrayOptionalCustom labels for scale points (e.g., ["Not confident", "Very confident"])
min_valuenumber0Minimum slider value (slider only)
max_valuenumber100Maximum slider value (slider only)
stepnumber1Increment between selectable slider values (slider only)
left_labelstring"Not confident"Label at the low end of the slider (slider only)
right_labelstring"Very confident"Label at the high end of the slider (slider only)
label_requirement.requiredbooleanfalseWhether the confidence rating must be completed before moving on

Without labels, a Likert scale uses Potato's own wording, from "Guessing" through "Somewhat confident", "Fairly confident", and "Confident" to "Certain", truncated to scale_points.

Examples

Likert Confidence Scale

Naming every point removes the guesswork about what a middle value means, which matters when several annotators have to use the scale the same way:

yaml
annotation_schemes:
  - annotation_type: radio
    name: toxicity
    description: Is this comment toxic?
    labels: ["Toxic", "Not Toxic"]
 
  - annotation_type: confidence
    name: toxicity_confidence
    description: How confident are you in your toxicity judgment?
    target_schema: toxicity
    scale_type: likert
    scale_points: 5
    labels: ["Not at all confident", "Slightly confident", "Moderately confident", "Very confident", "Extremely confident"]

Slider Confidence Scale

A slider records a finer-grained value than a 5-point scale, at the cost of a slower answer and values that are harder to compare across annotators. The 0 to 100 range below matches the wording of the instruction:

yaml
annotation_schemes:
  - annotation_type: radio
    name: stance
    description: What stance does the author take?
    labels: ["Support", "Oppose", "Neutral"]
 
  - annotation_type: confidence
    name: stance_confidence
    description: Rate your confidence from 0 (guessing) to 100 (certain).
    target_schema: stance
    scale_type: slider
    min_value: 0
    max_value: 100
    step: 1
    left_label: "Guessing"
    right_label: "Certain"

Required Confidence Rating

An optional confidence rating is the one annotators skip when they are moving quickly, which leaves the ratings you do collect biased toward careful items. Set label_requirement.required to true on the confidence scheme to block submission until the rating is answered:

yaml
annotation_schemes:
  - annotation_type: multiselect
    name: topics
    description: Select all topics that apply.
    labels: ["Politics", "Economy", "Health", "Education"]
 
  - annotation_type: confidence
    name: topics_confidence
    description: How confident are you in your topic selections?
    target_schema: topics
    scale_type: likert
    scale_points: 3
    labels: ["Low", "Medium", "High"]

Standalone Confidence (No Target)

Confidence annotations can also be used without a target schema for general self-assessment, such as recording domain familiarity once per item:

yaml
annotation_schemes:
  - annotation_type: confidence
    name: task_familiarity
    description: How familiar are you with this topic area?
    scale_type: likert
    scale_points: 5
    labels: ["Not familiar", "Slightly familiar", "Somewhat familiar", "Very familiar", "Expert"]

Output Format

json
{
  "toxicity_confidence": {
    "labels": {
      "confidence": 4
    }
  }
}

For Likert scales, values run from 1 to scale_points. For sliders, values run from min_value to max_value, and the slider starts at the midpoint of that range.

Best Practices

  1. Pair the rating with a target schema - a confidence rating is most useful when target_schema links it to a specific decision
  2. Use Likert for speed - discrete scales are faster to answer and easier to compare across annotators
  3. Use sliders for fine-grained measurement - when you need precise confidence values for downstream analysis
  4. Make confidence required - optional ratings get skipped, which biases the ratings you collect
  5. Analyze confidence patterns - low-confidence items are good candidates for adjudication or additional annotations

Further Reading

For implementation details, see the source documentation.