Skip to content

Rubric Evaluation

Build multi-criteria evaluation grids in Potato for LLM output assessment, essay grading, translation quality, and any structured rubric-based annotation task.

The rubric evaluation annotation schema provides a structured grid interface for scoring content across multiple criteria on a defined scale. Use it for LLM output evaluation, essay grading, translation quality assessment, or any task that needs structured multi-dimensional scoring.

Potato rubric evaluation grid scoring a response on accuracy, relevance, and fluencyRubric evaluation in Potato

Overview

The rubric evaluation schema presents:

  • A grid of criteria each with its own rating scale
  • Scale labels ranging from Poor to Excellent (customizable)
  • Optional overall score that summarizes across criteria
  • Descriptions for each criterion to guide annotators

Scoring several named dimensions in one pass is what separates a rubric from a single quality rating, and it is why the schema suits human evaluation of generative model output, where "good" decomposes into claims that disagree with each other.

Quick Start

Each entry in criteria becomes a row, and its description is the only place an annotator learns what the row means, so write one per criterion:

yaml
annotation_schemes:
  - annotation_type: rubric_eval
    name: response_quality
    description: Evaluate the quality of this AI-generated response.
    scale_points: 5
    criteria:
      - name: Accuracy
        description: Is the information factually correct?
      - name: Relevance
        description: Does the response address the question?
      - name: Fluency
        description: Is the response well-written and natural?

Configuration Options

FieldTypeDefaultDescription
annotation_typestringRequiredMust be "rubric_eval"
namestringRequiredUnique identifier for this schema
descriptionstringRequiredInstructions displayed to annotators
scale_pointsinteger5Number of points on the rating scale
scale_labelsarray["Poor", "Below Average", "Average", "Good", "Excellent"]Labels for each scale point
criteriaarrayRequiredList of criteria objects, each with name and optional description
show_overallbooleanfalseShow an additional overall score row below the criteria

Potato truncates scale_labels to scale_points, so a list longer than the scale loses its tail. Give the two settings matching lengths when you customize either.

Examples

LLM Output Evaluation

Five points give raters a neutral midpoint and room on either side of it, which suits open-ended model output where most responses are neither failures nor exemplars. show_overall adds a summary row, worth turning on when downstream analysis needs one number per response as well as the breakdown:

yaml
annotation_schemes:
  - annotation_type: rubric_eval
    name: llm_eval
    description: Rate the quality of this model-generated response.
    scale_points: 5
    scale_labels:
      - Poor
      - Fair
      - Average
      - Good
      - Excellent
    show_overall: true
    criteria:
      - name: Helpfulness
        description: Does the response provide useful and actionable information?
      - name: Accuracy
        description: Is the response factually correct and free of hallucinations?
      - name: Harmlessness
        description: Is the response free of harmful, biased, or inappropriate content?
      - name: Coherence
        description: Is the response logically structured and easy to follow?

Essay Grading

A four-point scale removes the midpoint, which forces each criterion onto one side of the pass line. Use it where the rubric is meant to produce a judgment rather than a distribution:

yaml
annotation_schemes:
  - annotation_type: rubric_eval
    name: essay_grade
    description: Grade this student essay using the rubric below.
    scale_points: 4
    scale_labels:
      - Below Expectations
      - Approaching
      - Meets Expectations
      - Exceeds Expectations
    criteria:
      - name: Thesis
        description: Is there a clear and arguable thesis statement?
      - name: Evidence
        description: Does the essay use relevant evidence to support claims?
      - name: Organization
        description: Is the essay logically organized with clear transitions?
      - name: Grammar
        description: Is the writing free of grammatical and spelling errors?

Translation Quality Assessment

Three points keep the raters' thresholds aligned when the criteria have a clear ceiling, since a translation is adequate or it is not:

yaml
annotation_schemes:
  - annotation_type: rubric_eval
    name: translation_quality
    description: Evaluate the quality of this machine translation.
    scale_points: 3
    scale_labels:
      - Unacceptable
      - Acceptable
      - Perfect
    criteria:
      - name: Adequacy
        description: Does the translation convey the same meaning as the source?
      - name: Fluency
        description: Does the translation read naturally in the target language?
      - name: Terminology
        description: Are domain-specific terms translated correctly?

Output Format

json
{
  "response_quality": {
    "labels": {
      "Accuracy": 4,
      "Relevance": 5,
      "Fluency": 3
    },
    "overall": 4
  }
}

Each criterion maps to its selected scale value (1-indexed). The overall field is included only when show_overall is true.

Best Practices

  1. Keep criteria independent - each criterion should measure a distinct dimension to avoid redundant scoring
  2. Write clear descriptions - annotators should know exactly what each criterion measures without ambiguity
  3. Use 3-5 scale points - fewer points reduce cognitive load; more than 7 points rarely improves reliability
  4. Provide anchor examples - in the description, mention what constitutes each end of the scale
  5. Enable overall score for aggregation - show_overall is useful when you need a single summary metric alongside detailed breakdowns

Further Reading

For implementation details, see the source documentation.