# Potato > Potato is a free, open-source data annotation tool from the University of Michigan. It supports 61 annotation types and 24 display types across text, images, audio, video, 3D point clouds, robot episodes, and AI-agent evaluation, with zero-code YAML configuration, AI-assisted labeling, and built-in quality control. Every documentation and blog page below is also available as clean Markdown — append `.md` to the URL (for example, https://www.potatoannotator.com/docs/getting-started/installation.md). ## Documentation ### Getting Started - [Quick Start](https://www.potatoannotator.com/docs/getting-started/quick-start.md): Get Potato running in under 5 minutes. Install via pip, write a YAML config, launch the server, and annotate your first dataset — no coding required. - [Installation](https://www.potatoannotator.com/docs/getting-started/installation.md): Install Potato 2.8 via pip on macOS, Windows, or Linux. Covers Python requirements, virtual environments, the vision extra, and why anyone on 2.7.1 or earlier should upgrade. - [What's New](https://www.potatoannotator.com/docs/getting-started/whats-new-v2.md): What's new in Potato v2.x. Version 2.8 adds computer vision, 3D point clouds, depth maps, robot episodes and world-model evaluation, plus chance-corrected agreement over geometry and time. 61 annotation types, 24 display types. - [Configuration Basics](https://www.potatoannotator.com/docs/getting-started/configuration-basics.md): Learn Potato's YAML configuration format — task settings, data file paths, annotation schemes, output formats, and user management essentials. ### Guides - [Guides](https://www.potatoannotator.com/docs/guides/overview.md): Tool-agnostic explainers on annotation practice: scheme design, inter-annotator agreement, crowdsourcing, and agent evaluation, each with a worked example. - [What Is Data Annotation?](https://www.potatoannotator.com/docs/guides/what-is-data-annotation.md): A plain-language introduction to data annotation, what it is, the main task types (classification, span labeling, ranking, free-text), and how to run an annotation project with Potato. - [Choosing an Annotation Scheme](https://www.potatoannotator.com/docs/guides/choosing-an-annotation-scheme.md): How to map your research question to the right Potato annotation type, radio, multiselect, span, likert, slider, pairwise, best-worst, multirate, rubric, and more. - [Designing Data Formats for Annotation](https://www.potatoannotator.com/docs/guides/data-formats-for-annotation.md): How to structure input data (JSON, JSONL, CSV) for an annotation project, what fields Potato expects, and how to plan for clean export to training pipelines. - [Writing Effective Annotation Guidelines](https://www.potatoannotator.com/docs/guides/writing-annotation-guidelines.md): How to write an annotation codebook that produces consistent labels, clear definitions, worked examples, edge-case rules, and a pilot-and-revise loop. - [Text Annotation](https://www.potatoannotator.com/docs/guides/text-annotation.md): A complete guide to text annotation, classification, multi-label tagging, rating, and free-text, and how to build each kind of text task in Potato with copy-paste config. - [Span Annotation](https://www.potatoannotator.com/docs/guides/span-annotation.md): A complete guide to span annotation, highlighting regions of text, overlapping and nested spans, label colors, BIO/IOB tagging, and building span tasks in Potato. - [Named Entity Recognition](https://www.potatoannotator.com/docs/guides/named-entity-recognition.md): What named entity recognition (NER) is, common label sets, and how to build an NER annotation task in Potato with colored span labels and tooltips. - [Coreference Resolution](https://www.potatoannotator.com/docs/guides/coreference-resolution.md): What coreference annotation is, how to group mentions into entity chains, and how to set up a coreference task in Potato. - [Relation and Event Extraction](https://www.potatoannotator.com/docs/guides/relation-and-event-extraction.md): How to annotate relations between entities and structured events with triggers and arguments, using span linking and event annotation in Potato. - [Entity Linking](https://www.potatoannotator.com/docs/guides/entity-linking.md): How to annotate entity linking, connecting mentions in text to entries in a knowledge base like Wikidata, and set up a linking task in Potato. - [How to Annotate Threaded Conversations](https://www.potatoannotator.com/docs/guides/annotating-threaded-conversations.md): Render reply structure from reply_to, annotate whole threads and individual comments at once, link spans across comments, and import a ConvoKit corpus both ways. - [Rating Scales and Likert Design](https://www.potatoannotator.com/docs/guides/rating-scales.md): How to design rating scales for annotation, Likert vs. sliders, how many points to use, avoiding acquiescence bias, and building rating tasks in Potato. - [Pairwise and Best–Worst Scaling](https://www.potatoannotator.com/docs/guides/pairwise-and-best-worst.md): When to use comparative judgments instead of ratings, pairwise comparison and best-worst scaling (MaxDiff), and how to set them up in Potato. - [Audio Annotation](https://www.potatoannotator.com/docs/guides/audio-annotation.md): A complete guide to audio annotation in Potato, classification, tagging, sound event detection on the waveform, transcription, quality (MOS) ratings, emotion, and speaker diarization. - [Video Annotation](https://www.potatoannotator.com/docs/guides/video-annotation.md): How to annotate video in Potato, frame-by-frame navigation, temporal segment labeling, per-frame classification, and tracking objects across frames. - [Image Annotation](https://www.potatoannotator.com/docs/guides/image-annotation.md): How to annotate images in Potato, classification, multi-label tagging, bounding boxes, polygons, and landmarks, and export to COCO/YOLO. - [How to Annotate Whisper Transcripts](https://www.potatoannotator.com/docs/guides/annotating-whisper-transcripts.md): How to take Whisper or WhisperX output and turn it into a running annotation project: which output file to keep, when you need diarization, how to assign speakers, and how to export with the timings intact. - [How to Annotate Image Segmentation Masks](https://www.potatoannotator.com/docs/guides/image-segmentation-annotation.md): Polygons, brush masks and click-to-segment, when each is the right tool, how to write boundary guidelines annotators can follow, and how to measure agreement on a mask. - [How to Annotate YouTube Subtitles](https://www.potatoannotator.com/docs/guides/annotating-youtube-subtitles.md): How to download captions with yt-dlp and annotate them in Potato, what auto-generated captions can and cannot support, how to handle the video, and what to check before redistributing media. - [How to Label Objects by Typing Their Name](https://www.potatoannotator.com/docs/guides/open-vocabulary-object-detection.md): Open-vocabulary detection lets an annotator type a phrase and get every match boxed. How Grounding DINO works, what the licence allows, and where the approach breaks down. - [How to Annotate Video Object Tracking](https://www.potatoannotator.com/docs/guides/video-object-tracking.md): Track an object through a clip with model-assisted mask propagation, handle occlusion honestly, and measure temporal agreement with a tolerance sweep. - [How to Annotate 3D Point Clouds](https://www.potatoannotator.com/docs/guides/point-cloud-annotation.md): Label lidar and photogrammetry data with oriented 3D cuboids. Formats, level of detail, why rotation should be a quaternion, and how to measure agreement on a 3D box. - [How to Annotate Depth Maps](https://www.potatoannotator.com/docs/guides/depth-map-annotation.md): Depth files do not carry their own unit, zero means no-return rather than close, and the interesting range is never the full range. What to get right before annotating depth. - [How to Annotate Robot Episodes](https://www.potatoannotator.com/docs/guides/robot-episode-annotation.md): Segment a robot demonstration into phases, judge the outcome, and draw a dense reward curve. Why partial is a real outcome and why the time-series lanes matter more than the video. - [Inter-Annotator Agreement Explained](https://www.potatoannotator.com/docs/guides/inter-annotator-agreement.md): A practical guide to inter-annotator agreement, percent agreement, Cohen's and Fleiss' kappa, and Krippendorff's alpha, when to use each, and how Potato reports them. - [Gold Standards and Attention Checks](https://www.potatoannotator.com/docs/guides/gold-standards-and-attention-checks.md): How to use gold-standard items and attention checks to catch low-quality annotators and keep a project calibrated, with Potato configuration. - [Adjudication and Resolving Disagreement](https://www.potatoannotator.com/docs/guides/adjudication-and-disagreement.md): What to do when annotators disagree, adjudication workflows, aggregation by majority vote, and statistical models like MACE that weight annotators by competence. - [How Many Annotators Do You Need?](https://www.potatoannotator.com/docs/guides/how-many-annotators.md): How to decide annotator count and overlap for an annotation project, balancing agreement, cost, and statistical confidence, with Potato overlap settings. - [Agreement for Spans and Structured Outputs](https://www.potatoannotator.com/docs/guides/agreement-for-spans.md): Why Cohen's and Fleiss' kappa break down for span, NER, and structured annotation, and what to use instead: F1-as-agreement, exact vs partial match, and Krippendorff's unitized alpha. - [Aggregating Crowd Labels: Beyond Majority Vote](https://www.potatoannotator.com/docs/guides/aggregating-crowd-labels.md): How to combine many noisy annotations into one label using annotator models like Dawid-Skene and MACE, when to trust them, and how Potato estimates competence and infers labels. - [Statistical Power and Sample Size for Annotation Studies](https://www.potatoannotator.com/docs/guides/statistical-power-annotation.md): How many items you need for a result to mean something, why that is a different question from how many annotators per item, and how to avoid underpowered, over-claimed annotation and evaluation studies. - [Documenting Datasets and Models: Datasheets, Data Statements, and Model Cards](https://www.potatoannotator.com/docs/guides/documenting-datasets-and-models.md): A reference to the three standard documentation frameworks for annotated data and the models built on it, what each covers, when to use which, and how reproducibility reporting ties them together. - [How to Measure Inter-Annotator Agreement on Bounding Boxes](https://www.potatoannotator.com/docs/guides/measuring-agreement-on-bounding-boxes.md): Raw IoU is not agreement. How to measure whether annotators agree on bounding boxes, polygons and masks with chance correction, and why detection, classification and localization need separate numbers. - [How to Tell Review from Rubber-Stamping in Pre-Labelled Data](https://www.potatoannotator.com/docs/guides/detecting-rubber-stamped-prelabels.md): When annotators accept model pre-labels without reading them, every quality measure improves, including agreement. Timing is the only signal that separates review from rubber-stamping. - [LLM and Vision Pre-Annotation](https://www.potatoannotator.com/docs/guides/llm-pre-annotation.md): How to speed up annotation with LLM pre-labeling and human verification, in-context learning, option highlighting, and vision pre-annotation, using Potato's AI support. - [Active Learning for Annotation](https://www.potatoannotator.com/docs/guides/active-learning.md): What active learning is, when it helps, and which query strategies Potato supports (uncertainty, diversity, BADGE, BALD), so you label fewer items for the same model quality. - [Collecting RLHF and Preference Data](https://www.potatoannotator.com/docs/guides/rlhf-preference-data.md): How to collect human preference data for RLHF and model alignment, pairwise comparisons, rubric scoring, and justifications, with Potato. - [How to Evaluate AI Agents](https://www.potatoannotator.com/docs/guides/evaluating-ai-agents.md): An overview of evaluating AI agents and LLMs with human annotation, trajectory, step, span, and comparison-level evaluation, and which Potato tool fits each. - [Annotating Agent Trajectories](https://www.potatoannotator.com/docs/guides/agent-trajectory-annotation.md): How to annotate AI agent trajectories step by step, error taxonomies, severity scoring, and trajectory-level success, using Potato's trajectory evaluation. - [Process Reward Models and Step-Level Labeling](https://www.potatoannotator.com/docs/guides/process-reward-models.md): How to collect process reward (PRM) data by labeling agent steps as correct or incorrect, first-error and per-step modes, with Potato. - [Evaluating Tool Use and Function Calling](https://www.potatoannotator.com/docs/guides/tool-use-evaluation.md): How to annotate and evaluate an agent's tool calls and function calling across trace formats (OpenAI, Anthropic, ReAct, LangChain) with Potato per-turn ratings. - [Web-Agent Evaluation](https://www.potatoannotator.com/docs/guides/web-agent-evaluation.md): How to evaluate web-browsing agents with screenshots and action overlays, per-step web action correctness, using Potato's web agent display. - [Coding-Agent Evaluation](https://www.potatoannotator.com/docs/guides/coding-agent-evaluation.md): How to evaluate coding agents, reviewing diffs, terminal output, and SWE-bench/Aider/Claude Code traces, with Potato's coding trace display. - [How to Evaluate Multi-Agent Systems](https://www.potatoannotator.com/docs/guides/evaluating-multi-agent-systems.md): A practical guide to evaluating multi-agent LLM systems, attributing failures to the responsible agent and handoff, reviewing the interaction graph, and scoring each agent and the team. - [Rubric-Based LLM Evaluation](https://www.potatoannotator.com/docs/guides/rubric-based-llm-evaluation.md): How to evaluate LLM outputs against multiple weighted criteria (MT-Bench style) using Potato's rubric evaluation type. - [Evaluating Computer-Use and Multimodal Agents](https://www.potatoannotator.com/docs/guides/evaluating-computer-use-agents.md): How to human-evaluate computer-use and GUI agents, plus voice, video, and document agents, judging each action and click, scoring turn-taking, and grounding events in time. - [Pairwise Model Comparison](https://www.potatoannotator.com/docs/guides/pairwise-model-comparison.md): How to compare two models or two responses head-to-head with human annotators, including multi-dimensional comparison and bias controls, using Potato. - [RAG Evaluation](https://www.potatoannotator.com/docs/guides/rag-evaluation.md): How to evaluate retrieval-augmented generation with human annotation, retrieval relevance, answer faithfulness, and citation spans, using Potato. - [Live Agent Evaluation](https://www.potatoannotator.com/docs/guides/live-agent-evaluation.md): How to evaluate an AI agent in real time, pause, send instructions, take over, rollback, and branch, using Potato's live agent display. - [Detecting Hallucinations with Span Annotation](https://www.potatoannotator.com/docs/guides/detecting-hallucinations.md): How to find and label hallucinations and factual errors in model output using span annotation and MQM-style error marking in Potato. - [Human Evaluation of Generated Text](https://www.potatoannotator.com/docs/guides/human-evaluation-generated-text.md): How to run a defensible human evaluation of LLM and NLG output: defining criteria precisely, choosing absolute vs pairwise ratings, powering the study, and reporting enough to reproduce it. - [How to Evaluate Generated Video and World Models](https://www.potatoannotator.com/docs/guides/world-model-evaluation.md): Ask annotators to mark the frame where the world stops making sense rather than to rate plausibility. Break-point annotation is checkable, localised, and produces a real agreement statistic. - [How to Evaluate VLM Grounding and Pointing](https://www.potatoannotator.com/docs/guides/vlm-grounding-evaluation.md): Score grounding by IoU at several thresholds, pointing by point-in-region hit rate, and ungroundedness as its own answer. Why a point scored like a box always fails. - [Running a Study on Prolific and MTurk](https://www.potatoannotator.com/docs/guides/crowdsourcing-prolific-mturk.md): How to run a crowdsourced annotation study with Potato on Prolific or Amazon Mechanical Turk, linking participants, quality control, and fair pay. - [Deploying an Annotation Server](https://www.potatoannotator.com/docs/guides/deploying-annotation-server.md): How to deploy a Potato annotation server for a real study, authentication options, HTTPS, and production setup with Docker or a reverse proxy. - [Exporting Annotations for Machine Learning](https://www.potatoannotator.com/docs/guides/exporting-annotations-for-ml.md): How to export Potato annotations into ML-ready formats, JSON/JSONL, CoNLL, Hugging Face Datasets, spaCy, COCO, and YOLO, and what each is for. - [Multilingual and Low-Resource Annotation](https://www.potatoannotator.com/docs/guides/multilingual-low-resource-annotation.md): Annotating in languages beyond English: the diversity gap, participatory methods with native speakers, and how to localize the Potato interface with right-to-left support, fonts, and translated labels. - [How to Run Annotation on an Air-Gapped Network](https://www.potatoannotator.com/docs/guides/air-gapped-annotation.md): What "offline" has to mean for an annotation tool, why an @import in a stylesheet defeats most air-gap audits, and how to get models onto a machine with no route out. - [How to Let a Coding Agent Write Your Annotation Config](https://www.potatoannotator.com/docs/guides/configs-your-coding-agent-can-check.md): The common failure when an LLM writes a config is inventing an option that does not exist. A JSON Schema plus a validator turns that from a runtime surprise into an editor error. - [AI Annotation Tools Compared: Open-Source and Paid](https://www.potatoannotator.com/docs/guides/annotation-tools-compared.md): Compares 14 annotation tools and platforms across text, vision, 3D and agent evaluation: which use AI to pre-label, and what each free tier really includes. ### Core Concepts - [Core Concepts](https://www.potatoannotator.com/docs/core-concepts/overview.md): The parts of a Potato config you edit on almost every project: data formats, annotation schemes, user management, UI configuration, and instance display. - [Data Formats](https://www.potatoannotator.com/docs/core-concepts/data-formats.md): Input formats for Potato: text, JSON, JSONL, CSV, images, audio, and video. Output to CoNLL, HuggingFace Datasets, spaCy, COCO, YOLO, Parquet, and more. - [Annotation Schemes](https://www.potatoannotator.com/docs/core-concepts/annotation-schemes.md): Define annotation schemas in YAML — radio, checkbox, span, likert, slider, text, and 55 more types. Configure labels, questions, required fields, and display options. - [User Management](https://www.potatoannotator.com/docs/core-concepts/user-management.md): Manage annotators in Potato — configure user lists, passwords, roles, task assignments, overlap settings, and OAuth or SSO authentication for your project. - [UI Configuration](https://www.potatoannotator.com/docs/core-concepts/ui-configuration.md): Customize Potato's annotation interface — header layout, sidebar sections, keyboard shortcuts, instance ordering, pagination, and display block positioning. - [Instance Display](https://www.potatoannotator.com/docs/core-concepts/instance-display.md): Configure the instance_display block to control how text, images, audio, and video content appears alongside annotation controls in your Potato task. ### Annotation Types - [Annotation Types](https://www.potatoannotator.com/docs/annotation-types/overview.md): Potato ships 61 annotation types. These pages cover the ones most projects use, from radio buttons and Likert scales to spans, media, and agent traces. - [Radio & Multiselect](https://www.potatoannotator.com/docs/annotation-types/radio-multiselect.md): Configure radio buttons for single-choice and checkboxes for multi-label annotation in Potato. Covers labels, keyboard shortcuts, required fields, and display options. - [Likert Scales](https://www.potatoannotator.com/docs/annotation-types/likert-scales.md): Configure Likert scales in Potato to measure attitudes, opinions, and quality ratings on 5, 7, or custom-point scales with labeled endpoints and required validation. - [Span Annotation](https://www.potatoannotator.com/docs/annotation-types/span-annotation.md): Highlight and label text spans in Potato for NER, relation extraction, and sentiment. Configure span types, colors, overlapping spans, and keyboard shortcuts. - [Text & Number Input](https://www.potatoannotator.com/docs/annotation-types/text-number-input.md): Collect free-text responses and numeric values in Potato. Configure text boxes, number fields, validation rules, placeholder text, and character limits. - [Pairwise Comparison](https://www.potatoannotator.com/docs/annotation-types/pairwise-comparison.md): Configure side-by-side comparisons in Potato for preference learning, A/B testing, and output quality assessment with randomized presentation order. - [Slider](https://www.potatoannotator.com/docs/annotation-types/slider.md): Configure a continuous numeric slider in Potato with custom min, max, step size, default value, and optional label display for granular rating and scoring tasks. - [Select (Dropdown)](https://www.potatoannotator.com/docs/annotation-types/select-dropdown.md): Configure a dropdown select field in Potato for single-choice annotation with long option lists — searchable, with optional default values and keyboard navigation. - [Multirate (Matrix Rating)](https://www.potatoannotator.com/docs/annotation-types/multirate.md): Configure a rating matrix in Potato where annotators score multiple items on the same scale simultaneously — useful for comparative evaluation and rubric-based tasks. - [Pure Display](https://www.potatoannotator.com/docs/annotation-types/pure-display.md): Add read-only display blocks to Potato annotation interfaces — formatted text, instructions, dividers, and contextual guidance that annotators cannot edit or interact with. - [Image Annotation](https://www.potatoannotator.com/docs/annotation-types/image-annotation.md): Annotate images in Potato with bounding boxes, polygons, freehand drawing, and landmark points. Includes zoom, pan, multi-label classification, and AI pre-labeling. - [Audio Annotation](https://www.potatoannotator.com/docs/annotation-types/audio-annotation.md): Segment audio files in Potato and assign labels to time regions. Displays an interactive waveform with playback controls, speed adjustment, and time-boundary marking. - [Video Annotation](https://www.potatoannotator.com/docs/annotation-types/video-annotation.md): Annotate video in Potato — classify temporal segments, mark keyframes, and track objects with frame-by-frame navigation and event labeling across the timeline. - [Span Linking](https://www.potatoannotator.com/docs/annotation-types/span-linking.md): Create typed directional relationships between two annotated text spans in Potato. Configure link types, colors, and bidirectional linking for relation extraction tasks. - [Dialogue Annotation](https://www.potatoannotator.com/docs/annotation-types/dialogue-annotation.md): Annotate conversations in Potato with per-turn ratings, span labels, and custom display. Supports multi-speaker dialogues, chat transcripts, and threaded replies. - [Event Annotation](https://www.potatoannotator.com/docs/annotation-types/event-annotation.md): Annotate N-ary event structures in Potato with event triggers, typed arguments, and entity roles for information extraction and ACE-style event detection tasks. - [Entity Linking](https://www.potatoannotator.com/docs/annotation-types/entity-linking.md): Link span annotations in Potato to external knowledge bases — Wikidata, UMLS, or custom entity APIs — with typeahead search and configurable confidence scores. - [Triage](https://www.potatoannotator.com/docs/annotation-types/triage.md): Build a rapid accept/reject/skip triage interface in Potato for fast data screening, quality filtering, and corpus cleaning — with keyboard shortcuts for speed. - [Conversation Trees](https://www.potatoannotator.com/docs/annotation-types/conversation-trees.md): Annotate branching conversation trees in Potato — rate dialogue nodes, select preferred paths, and capture hierarchical multi-turn evaluation preferences. - [Coreference Chains](https://www.potatoannotator.com/docs/annotation-types/coreference.md): Mark coreference chains in Potato by grouping text spans that refer to the same entity. Supports nested mentions, cross-sentence chains, and chain visualization. - [Best-Worst Scaling](https://www.potatoannotator.com/docs/annotation-types/best-worst-scaling.md): Efficient comparative annotation with Best-Worst Scaling in Potato — automatically generates comparison tuples and converts selections to continuous quality scores. - [Soft Label](https://www.potatoannotator.com/docs/annotation-types/soft-label.md): Configure soft label annotation in Potato for probability distribution allocation across categories using sliders that must sum to a fixed total. - [Confidence Annotation](https://www.potatoannotator.com/docs/annotation-types/confidence-annotation.md): Add confidence ratings paired with other annotations in Potato using Likert scales or sliders to capture annotator certainty. - [Constant Sum](https://www.potatoannotator.com/docs/annotation-types/constant-sum.md): Configure constant sum annotation in Potato for fixed-budget point allocation across categories using number inputs or sliders. - [Range Slider](https://www.potatoannotator.com/docs/annotation-types/range-slider.md): Configure dual-thumb range slider annotation in Potato for selecting a numeric range with customizable bounds, step size, and endpoint labels. - [Semantic Differential](https://www.potatoannotator.com/docs/annotation-types/semantic-differential.md): Configure semantic differential scales in Potato for measuring attitudes using bipolar adjective pairs with configurable scale points. - [Hierarchical Multiselect](https://www.potatoannotator.com/docs/annotation-types/hierarchical-multiselect.md): Build expandable tree-based taxonomy selectors in Potato for hierarchical classification, topic tagging, and multi-level category annotation tasks. - [Card Sort](https://www.potatoannotator.com/docs/annotation-types/card-sort.md): Build drag-and-drop card sorting interfaces in Potato for grouping, categorization, and information architecture tasks with closed or open sorting modes. - [Rubric Evaluation](https://www.potatoannotator.com/docs/annotation-types/rubric-eval.md): Build multi-criteria evaluation grids in Potato for LLM output assessment, essay grading, translation quality, and any structured rubric-based annotation task. - [Extractive QA](https://www.potatoannotator.com/docs/annotation-types/extractive-qa.md): Build SQuAD-style question answering interfaces in Potato for span-based answer extraction, reading comprehension tasks, and passage highlighting annotation. - [Error Span](https://www.potatoannotator.com/docs/annotation-types/error-span.md): Build MQM-style error annotation interfaces in Potato for translation quality evaluation, text error marking, and typed error span annotation with severity scoring. - [Transcript Formats](https://www.potatoannotator.com/docs/annotation-types/transcript-formats.md): Every transcript and subtitle format Potato reads: Whisper, WhisperX, Deepgram, AssemblyAI, AWS Transcribe, SRT, WebVTT, TTML, YouTube captions, CTM, Praat TextGrid, and ELAN EAF, plus sidecar files and the normalized turn model. ### Vision & Spatial - [Vision and Spatial Annotation](https://www.potatoannotator.com/docs/vision-spatial/overview.md): Potato annotates images, video, gigapixel scans, 3D point clouds, depth maps and robot episodes, and reports chance-corrected agreement over all of them. Overview of the vision and spatial surface. - [Interactive Segmentation](https://www.potatoannotator.com/docs/vision-spatial/segmentation.md): Click an object and get a mask, in the browser, with no GPU and no network call per click. How Potato runs MobileSAM through ONNX Runtime Web, and how to configure it. - [Open-Vocabulary Text Prompting](https://www.potatoannotator.com/docs/vision-spatial/text-prompting.md): Type a phrase and every match in the image comes back boxed. Grounding DINO under Apache-2.0, running in the browser with no GPU, and why the quantization was chosen by measurement. - [Geometry Primitives](https://www.potatoannotator.com/docs/vision-spatial/geometry-primitives.md): Boxes, polygons, polylines, ellipses, 2D cuboids, keypoint sets with skeletons, and instance-keyed brush masks, with V7-style keyboard conventions and a legacy profile. - [Deep Zoom and Gigapixel Images](https://www.potatoannotator.com/docs/vision-spatial/deep-zoom.md): Serve a tile pyramid as DZI and IIIF Image API 3.0, and paint brush masks at the source's full resolution with no texture-size ceiling. - [3D Point Clouds and Cuboids](https://www.potatoannotator.com/docs/vision-spatial/point-clouds.md): Annotate PCD, PLY, LAS and KITTI point clouds with oriented 3D cuboids, points, polylines and per-point segments. Octree level of detail, slab panels, and quaternion rotation that survives a round trip. - [Depth Maps](https://www.potatoannotator.com/docs/vision-spatial/depth-maps.md): Read 16-bit PNG, TIFF, NPY, PFM and EXR depth with percentile windowing, colormaps, a metres readout under the cursor, and unprojection into the 3D viewer. - [Camera Calibration and 2D Verification](https://www.potatoannotator.com/docs/vision-spatial/calibration.md): Project every 3D cuboid into each calibrated camera image, so annotators edit in 3D and verify in 2D. Why 3D labelling without this has no feedback loop. - [Media Ingest](https://www.potatoannotator.com/docs/vision-spatial/media-ingest.md): Annotate files a browser cannot display, including multi-page 16-bit TIFF, HEIC, camera RAW, ProRes and MKV, through a cached transcoding proxy. - [Video Tracking and Mask Propagation](https://www.potatoannotator.com/docs/vision-spatial/video-tracking.md): Draw a mask on one frame and SAM 2 follows the object through the clip, measured at 0.974 to 0.979 IoU per frame with no decay. Runs server-side by design. - [Robot Episode Annotation](https://www.potatoannotator.com/docs/vision-spatial/robot-episodes.md): Annotate robot demonstrations with N synchronized video streams and M time-series lanes on one timeline. Phase segmentation, outcome, dense reward, and LeRobot, RLDS and HDF5 import. - [World-Model and Generative-Video Evaluation](https://www.potatoannotator.com/docs/vision-spatial/world-model-evaluation.md): Show 2 to N generated videos frame-locked on one clock and ask annotators to mark the frame at which the world stops making sense, then tag which physical property broke. - [VLM Grounding, Pointing and Region Captioning](https://www.potatoannotator.com/docs/vision-spatial/vlm-grounding.md): Bind referring expressions to regions, score pointing by point-in-region hit rate rather than IoU, and count ungroundedness separately instead of as a miss. - [Model Zoo and Licences](https://www.potatoannotator.com/docs/vision-spatial/model-zoo.md): The seven models Potato can run, what each costs to install, which licence applies, and where it runs. Nothing is bundled and nothing downloads unasked. - [Computer Vision Formats](https://www.potatoannotator.com/docs/vision-spatial/cv-formats.md): 15 import and 29 export formats, 11 round-tripping. COCO with RLE and crowd regions, YOLO, Pascal VOC, CVAT, Darwin both ways, KITTI, MOT, DAVIS, Cityscapes and more. ### Measurement & Integrity - [Measurement and Integrity](https://www.potatoannotator.com/docs/measurement/overview.md): Potato reports chance-corrected agreement over geometry, time, 3D cuboids, captions and world-model break-points, and records how the annotations were made. Overview of the measurement surface. - [Agreement over Geometry](https://www.potatoannotator.com/docs/measurement/geometry-agreement.md): How Potato reports chance-corrected agreement for bounding boxes and polygons — detection, classification and localization, with σ and a KS statistic against an empirical chance baseline. - [Agreement over Time](https://www.potatoannotator.com/docs/measurement/temporal-agreement.md): Temporal IoU for segment boundaries, boundary α, and why break-point agreement is reported as a tolerance sweep rather than at a single threshold. - [Mask Consensus with STAPLE](https://www.potatoannotator.com/docs/measurement/mask-consensus-staple.md): Which mask should the dataset record, and who drew it well? STAPLE estimates a latent boundary with per-rater sensitivity and specificity, and beats majority vote when careful annotators are outnumbered. - [Drawing Telemetry](https://www.potatoannotator.com/docs/measurement/drawing-telemetry.md): How a box, polygon or mask was produced — time per shape, stroke dynamics, revision counts, and AI-suggestion accept latency. The only signal that separates review from rubber-stamping. - [Air-Gapped and Offline Deployment](https://www.potatoannotator.com/docs/measurement/air-gap.md): Every stylesheet, script, font and icon serves from your install, across all 14 templates including login and signup. Verified live at 62 requests, zero external. - [Machine-Readable Specs](https://www.potatoannotator.com/docs/measurement/machine-readable-specs.md): A JSON Schema for the config and an OpenAPI document for the API, both generated from the code and checked in CI, so a coding agent can configure Potato without inventing options. ### Features - [Features](https://www.potatoannotator.com/docs/features/overview.md): What Potato does around the annotation question itself: AI pre-labelling, quality control, multi-phase workflows, export formats, and process telemetry. - [AI Support](https://www.potatoannotator.com/docs/features/ai-support.md): Integrate OpenAI, Claude, Gemini, Ollama, HuggingFace, vLLM, and OpenRouter with Potato for label suggestions, pre-annotation, and AI-powered keyword highlighting. - [Active Learning](https://www.potatoannotator.com/docs/features/active-learning.md): Use 5 active learning strategies in Potato — uncertainty sampling, BADGE, BALD, diversity-based, and hybrid ensemble — to cut annotation cost by up to 50%. - [Agentic Annotation](https://www.potatoannotator.com/docs/features/agentic-annotation.md): Evaluate AI agents in Potato with 15 trace format converters, 5 display types, and pre-built schemas for tool-use, web-browsing, coding, and chat agents. Includes PRM and rubric evaluation modes. - [Audio Annotation](https://www.potatoannotator.com/docs/features/audio-annotation.md): Configure waveform-based audio annotation in Potato. Display interactive playback controls, set time markers, and combine with radio, likert, or text schemas. - [Solo Mode](https://www.potatoannotator.com/docs/features/solo-mode.md): Run a complete annotation pipeline solo in Potato — a 12-phase LLM-human workflow covering seeding, labeling, adjudication, and refinement without a full annotation team. - [Image Annotation](https://www.potatoannotator.com/docs/features/image-annotation.md): Configure image annotation in Potato with bounding boxes, polygons, and classification labels. Supports zoom, pan, multi-label, and AI-assisted pre-labeling. - [Training Phase](https://www.potatoannotator.com/docs/features/training-phase.md): Create training and qualification phases in Potato with practice items, gold-standard answers, and pass/fail thresholds to certify annotators before the main task. - [Admin Dashboard](https://www.potatoannotator.com/docs/features/admin-dashboard.md): Monitor annotation progress, manage annotator accounts, view per-annotator statistics, and configure settings in real time via Potato's admin dashboard. - [Multi-Phase Workflows](https://www.potatoannotator.com/docs/features/surveyflow.md): Build multi-stage annotation workflows in Potato — combine training phases, annotation tasks, and custom survey pages with consent forms and conditional branching. - [Export Formats](https://www.potatoannotator.com/docs/features/export-formats.md): Export Potato annotations to JSON, JSONL, CSV, CoNLL, HuggingFace Datasets, spaCy, COCO, YOLO, Pascal VOC, and Apache Parquet for downstream ML pipelines. - [Quality Control](https://www.potatoannotator.com/docs/features/quality-control.md): Ensure annotation quality in Potato with attention check items, gold-standard validation, configurable annotator overlap, and Krippendorff's Alpha reporting. - [Behavioral Tracking](https://www.potatoannotator.com/docs/features/behavioral-tracking.md): Log detailed annotator interactions in Potato — time-on-task, click patterns, edit history, and UI events — for quality analysis and behavioral research studies. - [Annotation History](https://www.potatoannotator.com/docs/features/annotation-history.md): Track every annotation action in Potato with timestamps, annotator IDs, and version history — enabling full audit trails, undo support, and change detection. - [ICL Labeling](https://www.potatoannotator.com/docs/features/icl-labeling.md): Use in-context learning in Potato to have an LLM pre-label instances, then route ambiguous ones for human verification — scaling annotation with AI in the loop. - [Productivity Features](https://www.potatoannotator.com/docs/features/productivity.md): Speed up annotation in Potato with keyboard shortcuts, hover tooltips, AI-powered keyword highlighting, and smart label suggestions for faster annotator throughput. - [Passwordless Login](https://www.potatoannotator.com/docs/features/passwordless-login.md): Enable username-only login in Potato without passwords — ideal for classroom demos, quick studies, and tasks using external crowdsourcing platform authentication. - [Task Assignment](https://www.potatoannotator.com/docs/features/task-assignment.md): Control instance distribution in Potato — sequential, random, uniform overlap, and per-annotator custom assignments to manage workload and coverage precisely. - [Category Assignment](https://www.potatoannotator.com/docs/features/category-assignment.md): Route annotation instances to qualified annotators based on demonstrated expertise in Potato. Configure category-based assignment rules and role-based filtering. - [Visual AI Support](https://www.potatoannotator.com/docs/features/visual-ai-support.md): Use vision LLMs — GPT-4 Vision, Claude Vision, Gemini, and YOLO — to pre-annotate images, generate bounding box suggestions, and assist with visual tasks in Potato. - [Layout Customization](https://www.potatoannotator.com/docs/features/layout-customization.md): Design custom annotation layouts in Potato with HTML templates and CSS — side-by-side comparisons, multi-column schemas, and fully custom display block structures. - [MACE Competence Estimation](https://www.potatoannotator.com/docs/features/mace.md): Use the MACE algorithm in Potato to estimate annotator competence, weight disagreements, and infer gold-standard labels from noisy multi-annotator annotation data. - [Option Highlighting](https://www.potatoannotator.com/docs/features/option-highlighting.md): AI-powered option highlighting in Potato pre-selects likely correct labels using LLMs, reducing cognitive load and accelerating throughput on discrete annotation tasks. - [Diversity Ordering](https://www.potatoannotator.com/docs/features/diversity-ordering.md): Reorder annotation instances in Potato using embedding-based diversity scoring to maximize dataset coverage and reduce redundancy in large unlabeled corpora. - [Remote Data Sources](https://www.potatoannotator.com/docs/features/remote-data-sources.md): Load annotation data dynamically in Potato from HTTP URLs, S3 buckets, Google Cloud Storage, PostgreSQL databases, and HuggingFace Datasets without local files. - [Survey Instruments](https://www.potatoannotator.com/docs/features/survey-instruments.md): 55 validated survey questionnaires built into Potato — Big Five personality, PHQ-9, GAD-7 mental health, PANAS affect, and demographic scales, ready to embed. - [Parquet Export](https://www.potatoannotator.com/docs/features/parquet-export.md): Export Potato annotations to Apache Parquet — a columnar format optimized for large-scale ML pipelines, Spark, DuckDB, Pandas, and HuggingFace Datasets integration. - [Web Agent Annotation](https://www.potatoannotator.com/docs/features/web-agent-annotation.md): Review web-browsing agent traces in Potato with filmstrip navigation, SVG overlays (clicks, bounding boxes, mouse paths), and per-step annotation controls. - [Live Agent Evaluation](https://www.potatoannotator.com/docs/features/live-agent-evaluation.md): Watch AI agents work in real time and annotate their behavior mid-execution with pause, instruct, and takeover controls. Supports web and coding agents with Anthropic, Ollama, and Claude SDK. - [Coding Agent Annotation](https://www.potatoannotator.com/docs/features/coding-agent-annotation.md): Annotate coding agent traces with diff rendering, terminal output, and file tree navigation. Import from Claude Code, Aider, SWE-Agent, and other coding assistants. - [Webhooks](https://www.potatoannotator.com/docs/features/webhooks.md): Send HMAC-signed HTTP webhook notifications from Potato for 7 annotation event types — with exponential backoff retry, admin monitoring, and Python/Node verification. - [Process Reward Annotation](https://www.potatoannotator.com/docs/features/process-reward-annotation.md): Collect per-step reward signals for training process reward models with first-error and per-step annotation modes. Export directly to PRM, DPO, and SWE-bench training formats. - [Code Review Annotation](https://www.potatoannotator.com/docs/features/code-review-annotation.md): Review AI coding agent output with GitHub PR-style inline diff comments, file-level correctness ratings, and approve or reject verdicts for code quality evaluation. - [Live Coding Agent Observation](https://www.potatoannotator.com/docs/features/live-coding-agent.md): Watch coding agents work in real time with pause, rollback, and branching. Three backends supported: Ollama for local models, Anthropic API, and Claude Agent SDK. - [Psychometrics Engine](https://www.potatoannotator.com/docs/features/psychometrics.md): Item response theory for annotation. Every label gets a posterior probability and a confidence interval instead of a bare majority vote, and the model needs neither gold labels nor an LLM to fit. - [Multiplayer Rooms](https://www.potatoannotator.com/docs/features/multiplayer-rooms.md): Run live norming and calibration sessions inside Potato. Blind votes, a host reveal, discussion, and a real-time Krippendorff's alpha meter that shows what the session was worth. - [Boundary Lab](https://www.potatoannotator.com/docs/features/boundary-lab.md): Probe the decision boundary, not just the point. After each label, Potato asks whether a minimal counterfactual edit would change it — producing contrast sets and quality control from ordinary annotation. - [Truth Serum](https://www.potatoannotator.com/docs/features/truth-serum.md): Surprisingly-popular scoring for annotation. One extra micro-question per label lets Potato beat majority vote on hard items, with no gold labels anywhere. - [Think-Aloud Mode](https://www.potatoannotator.com/docs/features/think-aloud.md): Annotators talk while they work and Potato stores the verbatim transcript as the rationale. Speech-to-text runs fully locally with faster-whisper, so no audio leaves the machine and there is no cloud API or LLM in the pipeline. - [Paper Mode](https://www.potatoannotator.com/docs/features/paper-mode.md): Turn an annotation project into a compilable LaTeX dataset report with one command. Label distributions, agreement statistics with correct citations, annotator tables, timing, and limitations. - [Pocket Mode](https://www.potatoannotator.com/docs/features/pocket-mode.md): Annotate from a phone or tablet. Touch devices are routed automatically to a swipeable card stack with thumb-zone controls, offline annotation with sync, and home-screen install as a PWA. - [Keystroke Logging](https://www.potatoannotator.com/docs/features/keystroke-logging.md): Potato can record the pauses, bursts, revisions and pastes behind a free-text answer without recording any of the characters an annotator types. - [Writing-Process Detection](https://www.potatoannotator.com/docs/features/writing-process-detection.md): Turn Potato's keystroke logs into auditable flags for pasted, transcribed, or machine-generated free-text answers, with thresholds you can read and override. - [Keystroke Logging Ethics](https://www.potatoannotator.com/docs/features/keystroke-logging-ethics.md): Consent, IRB review, retention, and participant rights for Potato's keystroke logging: what the data can reveal and how to use writing-process flags fairly. ### Agent Evaluation - [Agent Evaluation](https://www.potatoannotator.com/docs/agent-evaluation/overview.md): Human and LLM-judge evaluation of agent output: judge calibration, trace review, trajectory editing, programmatic evaluators, and CI evaluation. - [LLM-as-Judge Calibration](https://www.potatoannotator.com/docs/agent-evaluation/judge-calibration.md): Auto-label data with one or more LLM judges, then run a blind human calibration pass to measure accuracy, agreement, and calibration error. Answers "should I trust this LLM judge?" with a defensible, reproducible workflow. - [Judge ↔ Human Alignment](https://www.potatoannotator.com/docs/agent-evaluation/judge-alignment.md): Measure how well an LLM judge agrees with your human gold labels. Potato runs the judge over annotated instances, computes Cohen's kappa, a confusion matrix, and a disagreement list, and tracks agreement as you refine the rubric. - [Signal-Based Triage Queue](https://www.potatoannotator.com/docs/agent-evaluation/triage-queue.md): Prioritize the annotation queue by a per-item quality signal so reviewers see the worst or most-suspect traces first, instead of annotating in arrival order. Route by agent errors, production thumbs-down, low scores, or any custom field. - [Trajectory Editing for SFT/DPO](https://www.potatoannotator.com/docs/agent-evaluation/trajectory-correction.md): Annotators rewrite the steps of an agent trace to fix a wrong reasoning step, correct a tool call, or strengthen the final answer, and Potato exports each original/corrected pair as supervised fine-tuning targets and DPO preference pairs. - [Three-Pane Trace Evaluation (eval_trace)](https://www.potatoannotator.com/docs/agent-evaluation/eval-trace.md): The eval_trace display splits one agent trace into three synchronized panes (Reasoning, Function Calls, and Final Answer) so an evaluator sees what the agent thought, did, and produced at a glance. Built for continuous evaluation. - [Programmatic Evaluators](https://www.potatoannotator.com/docs/agent-evaluation/programmatic-evaluators.md): Score agent trajectories and text outputs automatically with Potato's Flask-free evaluator library — deterministic trajectory match, tool-use correctness, reference-free LLM-as-judge, and heuristics (exact match, edit distance, JSON, embeddings). - [Datasets & Experiments](https://www.potatoannotator.com/docs/agent-evaluation/datasets-and-experiments.md): Build versioned evaluation datasets and run experiments that score agent outputs over time. Potato's eval backbone — file or SQLite storage, tagged versions, splits, SFT/DPO export, and a side-by-side experiment comparison with regression deltas. - [Automation Rules](https://www.potatoannotator.com/docs/agent-evaluation/automation-rules.md): Route incoming agent traces automatically with filter, sample, and action rules. Potato runs each rule over every incoming trace to route it to the annotation queue, curate it into a dataset, run an evaluator, fire a webhook, or notify annotators. - [CI Evaluation](https://www.potatoannotator.com/docs/agent-evaluation/ci-evaluation.md): Run Potato evaluations inside your own pytest suite and gate CI on score thresholds, so a prompt or model change that regresses agent quality fails the build like a unit test. Includes an expect() assertion API and an example GitHub Actions workflow. - [Model Arena](https://www.potatoannotator.com/docs/agent-evaluation/model-arena.md): Send one prompt to N models side by side, compare their responses, and pick the best to build a win-rate leaderboard. Provider-agnostic — compare OpenAI, Anthropic, Ollama, vLLM, and Gemini in one view, free and self-hosted. - [Semantic Curation (Catalog)](https://www.potatoannotator.com/docs/agent-evaluation/semantic-curation.md): Find which agent traces to review by similarity, not just rules. An embedding index over your items powers similarity search ("find traces like this failure") and dynamic slices, which are saved semantic and metadata filters that auto-include new matching traces and curate them into datasets. - [Tracing SDK (potato_trace)](https://www.potatoannotator.com/docs/agent-evaluation/tracing-sdk.md): Instrument any agent with the lightweight potato_trace SDK to capture its runs into Potato for evaluation. Decorate functions with @traceable (sync or async) and nested runs are captured and sent to Potato's ingestion webhook, with optional OpenTelemetry export. - [Multi-Agent Team Evaluation](https://www.potatoannotator.com/docs/agent-evaluation/multi-agent-evaluation.md): Annotate multi-agent systems by team structure, not a flat transcript. Potato adds a clickable agent-interaction graph, cross-agent failure attribution, handoff review, per-agent and per-team scorecards, a tool-contention timeline, and emergent-behavior tagging. - [Multimodal-Agent Evaluation](https://www.potatoannotator.com/docs/agent-evaluation/multimodal-agent-evaluation.md): Evaluate agents that act beyond text, computer-use and GUI agents, voice assistants, video, and document agents. Potato adds purpose-built schemas for GUI trajectories with click grounding, full-duplex voice timelines, video temporal grounding with live IoU, speech-transcript error tagging, interleaved multimodal reasoning, and table-grid structure. ### Qualitative Data Analysis - [QDA Mode](https://www.potatoannotator.com/docs/qda/qda-mode.md): Turn Potato into a collaborative qualitative data analysis workspace. QDA Mode composes a living codebook, in-vivo coding, analyst memos, cases, and full-text search for coding interview transcripts, open-ended survey responses, and field notes. ### Tools & Utilities - [Tools & Utilities](https://www.potatoannotator.com/docs/tools/overview.md): Command-line utilities that ship with Potato: config preview, migration between versions, a synthetic user simulator, and the debugging guide. - [Preview CLI](https://www.potatoannotator.com/docs/tools/preview-cli.md): Use Potato's preview CLI to validate YAML configs, inspect annotation schemas, and render a static preview of your interface without starting the full server. - [Migration CLI](https://www.potatoannotator.com/docs/tools/migration-cli.md): Migrate Potato configuration files from v1 to v2 format automatically. Includes a dry-run mode to preview all changes before writing them to disk. - [User Simulator](https://www.potatoannotator.com/docs/tools/simulator.md): Simulate multiple concurrent annotators in Potato for integration testing — configure annotation strategies, speed, and agreement levels for realistic load tests. - [Debugging Guide](https://www.potatoannotator.com/docs/tools/debugging-guide.md): Debug Potato annotation projects — enable verbose logging, use debug flags, interpret common error messages, and troubleshoot data loading and schema validation issues. ### Deployment - [Deployment](https://www.potatoannotator.com/docs/deployment/overview.md): Running Potato for real annotators: local development, production setup, authentication, task assignment, and crowdsourcing platform integration. - [Local Development](https://www.potatoannotator.com/docs/deployment/local-development.md): Install and run Potato locally for development and testing. Covers pip install, config setup, launching the dev server, and debugging common startup issues. - [Production Setup](https://www.potatoannotator.com/docs/deployment/production-setup.md): Deploy Potato in production with nginx, SSL/TLS, gunicorn, systemd, and Docker. Covers reverse proxy configuration, HTTPS setup, and multi-user performance tuning. - [Crowdsourcing Integration](https://www.potatoannotator.com/docs/deployment/crowdsourcing.md): Integrate Potato with Prolific and Amazon MTurk for crowdsourced annotation. Covers completion URLs, participant ID tracking, attention checks, and payment configuration. - [MTurk Integration](https://www.potatoannotator.com/docs/deployment/mturk-integration.md): Run Potato annotation tasks on Amazon Mechanical Turk — configure HITs, qualification tests, approval workflows, bonus payments, and annotator quality filtering. - [Data Directory Loading](https://www.potatoannotator.com/docs/deployment/data-directory.md): Configure Potato to load annotation instances from a folder — supports glob patterns, live watching for new files, and recursive subdirectory scanning with filters. - [SSO & OAuth Authentication](https://www.potatoannotator.com/docs/deployment/sso-oauth.md): Configure Google OAuth, GitHub OAuth, and generic OIDC in Potato. Restrict access by email domain or GitHub org, and enable mixed-mode login with passwords. - [Password Management](https://www.potatoannotator.com/docs/deployment/password-management.md): Configure PBKDF2-SHA256 password hashing, admin CLI/API resets, self-service token-based reset flows, and SQLite or PostgreSQL credential storage in Potato. - [Task Assignment](https://www.potatoannotator.com/docs/deployment/task-assignment.md): Control how Potato distributes annotation items to annotators. Covers all assignment strategies including the custom Batch strategy for repeat-round studies, and reclaiming abandoned assignments from Prolific or QC-blocked workers. - [Heterogeneous Annotator Coverage](https://www.potatoannotator.com/docs/deployment/heterogeneous-coverage.md): Assign different numbers of annotators to different items. Configure a default cap, a stratified overlap sample for quality monitoring, adaptive disagreement boosts, per-annotator quotas, and automatic adjudication routing. - [Reverse Proxy (URL Path Prefix)](https://www.potatoannotator.com/docs/deployment/reverse-proxy.md): Serve Potato under a sub-path behind a reverse proxy, such as https://host/app1/. Configure a deployment URL prefix so static assets, annotation actions, and live streams resolve correctly under the mount path. ### API Reference - [API Overview](https://www.potatoannotator.com/docs/api/overview.md): HTTP API reference for Potato — admin endpoints for managing annotators, triggering exports, querying progress statistics, and integrating with external pipelines. ### Contributing - [Contributing](https://www.potatoannotator.com/docs/contributing/overview.md): Contribute to Potato — set up a dev environment, run the test suite, submit pull requests, report bugs, and write documentation for the open-source annotation tool. ## Showcase designs - [Acoustic Scene Classification](https://www.potatoannotator.com/showcase/acoustic-scene-classification): Classify audio recordings by acoustic environment following the TUT/DCASE dataset format. - [Annotator Demographics with Consent](https://www.potatoannotator.com/showcase/annotator-demographics-consent): A subjective offensiveness-rating task wrapped in an informed-consent page and a standardized demographic survey, so you can analyze labels by annotator background. - [Audio-Visual Sentiment Analysis](https://www.potatoannotator.com/showcase/audio-sentiment-analysis): Rate sentiment in speech segments following CMU-MOSI and CMU-MOSEI multimodal annotation protocols. - [Audio Transcription Review](https://www.potatoannotator.com/showcase/audio-transcription): Review and correct automatic speech recognition transcripts with waveform display. - [AudioSet Event Classification](https://www.potatoannotator.com/showcase/audioset-event-classification): Multi-label audio event tagging following the AudioSet ontology for weak supervision. - [Coreference Chains](https://www.potatoannotator.com/showcase/coreference): Group coreferring text mentions into chains with visual highlighting. Combines span annotation for mention detection with coreference grouping. - [Conversation Tree](https://www.potatoannotator.com/showcase/conversation-tree): Hierarchical conversation tree annotation with per-node quality ratings and path selection. Ideal for evaluating chatbot responses and dialogue systems. - [Continuous Emotion Rating](https://www.potatoannotator.com/showcase/emotion-dimensional-rating): Rate emotional dimensions (valence, arousal, dominance) continuously following MSP-IMPROV protocol. - [Entity Linking](https://www.potatoannotator.com/showcase/entity-linking): Span annotation with knowledge base linking to Wikidata. Annotators highlight entities and link them to their corresponding Wikidata entries via an inline search widget. - [Environmental Sound Classification](https://www.potatoannotator.com/showcase/environmental-sound-classification): Classify environmental sounds into categories following UrbanSound8K and ESC-50 datasets. - [Event Annotation](https://www.potatoannotator.com/showcase/event-annotation): N-ary event annotation with trigger spans and typed argument roles. Annotate events like ATTACK, HIRE, and TRAVEL with constrained entity arguments and hub-spoke arc visualization. - [Image Classification](https://www.potatoannotator.com/showcase/image-classification): Multi-label image classification with quality assessment for computer vision datasets. - [Keyword Spotting](https://www.potatoannotator.com/showcase/keyword-spotting): Classify spoken commands and keywords following the Google Speech Commands dataset format. - [Music Tagging](https://www.potatoannotator.com/showcase/music-tagging): Multi-label music tagging following MagnaTagATune dataset format for instrument and genre annotation. - [Named Entity Recognition](https://www.potatoannotator.com/showcase/named-entity-recognition): Span-based annotation for identifying entities like persons, organizations, locations, and dates in text. - [LLM Response Preference](https://www.potatoannotator.com/showcase/pairwise-preference): Compare AI-generated responses to collect preference data for RLHF training. - [Respiratory Sound Classification](https://www.potatoannotator.com/showcase/respiratory-sound-classification): Classify lung and respiratory sounds for medical diagnosis following ICBHI 2017 Challenge format. - [Sentiment Analysis](https://www.potatoannotator.com/showcase/sentiment-analysis): Simple 3-way sentiment classification with radio buttons. Perfect for social media analysis, product reviews, and customer feedback. - [Sound Event Detection](https://www.potatoannotator.com/showcase/sound-event-detection): Temporal sound event annotation with strong labels following DCASE Challenge protocols. - [Speaker Diarization](https://www.potatoannotator.com/showcase/speaker-diarization): Segment and label speakers in multi-party conversations following AMI Meeting Corpus guidelines. - [Speech Emotion Recognition](https://www.potatoannotator.com/showcase/speech-emotion-recognition): Classify emotional content in speech following IEMOCAP and CREMA-D annotation schemes. - [Speech Intelligibility Rating](https://www.potatoannotator.com/showcase/speech-intelligibility-rating): Rate speech intelligibility for pathological speech following TORGO database annotation protocols. - [Speech Quality MOS Rating](https://www.potatoannotator.com/showcase/speech-quality-mos): Rate speech quality using Mean Opinion Score following ITU-T P.800 and Blizzard Challenge protocols. - [User Feedback Survey](https://www.potatoannotator.com/showcase/survey-feedback): Comprehensive survey template for collecting user feedback with Likert scales and open-ended questions. - [Toxicity Detection](https://www.potatoannotator.com/showcase/toxicity-detection): Multi-label classification for identifying various types of toxic content including hate speech, threats, and harassment. - [Transcript Format Ingestion](https://www.potatoannotator.com/showcase/transcript-formats): Six transcript formats rendered as the same speaker bubbles, with per-turn labels and spans that work regardless of where the transcript came from. - [Triage](https://www.potatoannotator.com/showcase/triage): Rapid accept/reject/skip screening interface for high-throughput data quality filtering. Auto-advances to the next item after each decision. ## Statistics tools - [Inter-annotator agreement calculator](https://www.potatoannotator.com/tools/inter-annotator-agreement): Krippendorff's alpha (nominal/ordinal/interval), Cohen's kappa, Fleiss' kappa, and bootstrap confidence intervals from a CSV, computed in the browser. - [Dawid-Skene consensus calculator](https://www.potatoannotator.com/tools/dawid-skene): consensus labels and per-annotator reliability from noisy crowd labels via EM. - [Bradley-Terry ranking calculator](https://www.potatoannotator.com/tools/bradley-terry): strengths, win probabilities, and Elo from pairwise preference judgments. - [Annotation power planner](https://www.potatoannotator.com/tools/annotation-power): Monte-Carlo estimate of how many items you need for a target confidence-interval width on Krippendorff's alpha. ## Blog - [Agreement Cannot Catch Rubber-Stamping](https://www.potatoannotator.com/blog/agreement-cannot-catch-rubber-stamping.md): When annotators accept model pre-labels without reading them, inter-annotator agreement goes up. Every quality measure computed from the annotations improves. Timing is the only signal left. - [Partial Is a Real Outcome: Annotating Robot Episodes](https://www.potatoannotator.com/blog/partial-is-a-real-outcome.md): Robot demonstrations are mostly partial successes, and a binary outcome throws that away. What a timeline-shaped annotation interface has to get right, from min/max downsampling to hindsight relabelling. - [Shipping Segmentation That Runs in a Browser Tab](https://www.potatoannotator.com/blog/segmentation-in-the-browser.md): Click-to-segment and open-vocabulary text prompting both run client-side with no GPU. What it took: a verified encoder contract, a quantization chosen by measurement, and a tokenizer written by hand. - [Raw IoU Is Not Agreement](https://www.potatoannotator.com/blog/raw-iou-is-not-agreement.md): Mean IoU between annotators reads about 0.95 on a typical detection corpus, including for annotators who never looked at the image. What a chance-corrected alternative looks like, and why Krippendorff's alpha over IoU distance does not work. - [Potato 2.8: Annotate Anything, Then Measure It](https://www.potatoannotator.com/blog/potato-2-8-release.md): Potato 2.8 adds computer vision, gigapixel deep zoom, 3D point clouds, depth maps, robot episodes, generative-video evaluation and VLM grounding — and reports chance-corrected agreement over all of them. - [Reading the Writing Process: Keystroke Logging for Free-Text Annotation](https://www.potatoannotator.com/blog/keystroke-logging-writing-process.md): Potato can now record how annotators produce free-text answers, without recording what they type, and turn the pauses, revisions and pastes into auditable flags for responses that were pasted rather than written. - [Annotating ASR Transcripts: A Worked Example](https://www.potatoannotator.com/blog/annotating-asr-transcripts.md): A full walkthrough from a folder of Whisper output to labeled speaker turns: choosing an annotation unit, handling diarization, writing the config, running the task, and exporting with the time alignment intact. - [Judging the Judge: Human and Model Reasoning, Side by Side](https://www.potatoannotator.com/blog/human-and-model-chain-of-thought.md): An LLM judge is only as trustworthy as the human labels you validated it against. Put a person's chain of thought next to a model's on the same item, and check your judge against labels that carry their own confidence intervals. - [Potato 2.7.1: The Transcript Already Exists](https://www.potatoannotator.com/blog/potato-2-7-1-transcripts.md): Potato 2.7.1 reads 21 transcript and subtitle formats directly, loads them from sidecar files next to your media, and ships a converter that turns a folder of ASR output into an annotation-ready data file. - [Instrumenting the Calibration Meeting](https://www.potatoannotator.com/blog/norming-sessions-and-think-aloud.md): Annotation teams already run norming sessions and already wish they knew how their annotators think. Multiplayer Rooms measure the calibration meeting with a live agreement meter; Think-Aloud Mode records the reasoning out loud, locally, with no LLM. - [Quality Control Without Gold Labels](https://www.potatoannotator.com/blog/quality-control-without-gold-labels.md): Gold standards are expensive, leaky, and only test the items you thought to plant. Two alternatives, peer-prediction scoring and counterfactual boundary probes, catch inattentive annotators and broken guidelines without a single planted item. - [Labels With Error Bars: Item Response Theory for Annotation](https://www.potatoannotator.com/blog/labels-with-error-bars.md): Majority vote throws away most of what your annotators told you. Item response theory gives every label a posterior probability and a confidence interval, estimates annotator ability and item difficulty, and finds broken codebook entries, with no gold labels and no LLM. - [Potato 2.7: Measuring How Labels Are Made](https://www.potatoannotator.com/blog/potato-2-7-release.md): Potato 2.7 adds labels with error bars, live norming rooms, counterfactual boundary probes, peer-prediction scoring, local voice rationales, one-command dataset reports, and mobile annotation, none of which need an LLM. Plus multi-agent and multimodal agent evaluation. - [Validated Survey Instruments for Annotation Studies: Personality, Affect, Wellbeing, and Demographics](https://www.potatoannotator.com/blog/survey-instruments-for-annotation-studies.md): When who annotates matters, a validated questionnaire beats a question you invented. A tour of Potato's 55 built-in survey instruments and when each one belongs in your study. - [Documenting Your Annotation Dataset: Data Statements, Datasheets, and a Release Checklist](https://www.potatoannotator.com/blog/documenting-annotation-datasets.md): What to record when you release an annotation dataset: the curation rationale, the annotator pool, the guidelines, and the intended use, plus how much of it Potato captures for you. - [Potato 2.0 at ACL 2026](https://www.potatoannotator.com/blog/potato-2-acl-2026.md): Our paper on Potato 2.0 is in the ACL 2026 System Demonstrations. Here is what it covers and how to cite Potato in your work. - [Evaluating Voice and Video Agents](https://www.potatoannotator.com/blog/evaluating-voice-and-video-agents.md): A walkthrough of human evaluation for spoken, video, and document agents in Potato: scoring turn-taking on a dual-track timeline, grounding video events with live IoU, tagging speech errors, and marking table structure. - [Evaluating Computer-Use Agents, Step by Step](https://www.potatoannotator.com/blog/computer-use-agent-evaluation.md): A walkthrough of human evaluation for computer-use and GUI agents in Potato: judging each action, checking click grounding on the screenshot, and reviewing tool calls one at a time. - [Debugging Multi-Agent Failures: A Walkthrough](https://www.potatoannotator.com/blog/debugging-multi-agent-systems.md): How to find why a multi-agent LLM system failed using Potato: the interaction graph, failure attribution, handoff review, per-agent scorecards, tool-contention timeline, and emergent-behavior tagging. - [Potato 2.6: Qualitative Data Analysis Meets Agent Evaluation](https://www.potatoannotator.com/blog/potato-2-6-release.md): Potato 2.6 is out: QDA Mode for qualitative coding, an LLM-as-judge calibration and alignment workflow, trajectory editing that produces SFT and DPO training data, a 3x faster boot, and a relicense to GPL-3.0-or-later. - [Potato 2.6.2: A Complete Open-Source Agent-Evaluation Suite](https://www.potatoannotator.com/blog/potato-2-6-2-agent-evaluation-suite.md): The 2.6.x line turns Potato into a full, free agent-evaluation platform: trace ingestion from OpenTelemetry, LangGraph, CrewAI, and AutoGen, multi-agent team annotation with a clickable interaction graph, multimodal-agent schemas for GUI, voice, and video, plus a model arena, CI gating, and curation. - [Disagreement Is Signal, Not Noise: When to Keep Annotator Disagreement Instead of Resolving It](https://www.potatoannotator.com/blog/disagreement-is-signal-not-noise.md): Annotation pipelines are built to erase disagreement, but on subjective tasks the disagreement is the data. A guide to telling genuine label variation from error, and keeping it in Potato. - [Adaptive Annotator Coverage for Large Datasets](https://www.potatoannotator.com/blog/adaptive-annotator-coverage.md): Potato 2.6 lets you assign one annotator to most items and three to a stratified sample, raise coverage on items where annotators disagree, and route the contested ones to an adjudicator. - [Collecting Annotator Demographics Responsibly: What to Ask and How to Ask It](https://www.potatoannotator.com/blog/collecting-annotator-demographics-responsibly.md): Annotator identity shapes labels on subjective tasks, so demographics are worth collecting, but only with consent and a reason. What to ask, what to leave alone, and how to run the flow in Potato. - [Closing the Loop: Routing Agent Errors and Judge Disagreements Back to Humans](https://www.potatoannotator.com/blog/closing-the-loop-judge-alignment-triage.md): Human review time is the scarcest resource in agent evaluation. Potato 2.6 pairs a signal-based triage queue with judge-human alignment so the worst traces reach people first and your LLM judge keeps getting better. - [From Evaluation to Training Data: Trajectory Editing for SFT and DPO](https://www.potatoannotator.com/blog/trajectory-editing-sft-dpo-training-data.md): Most agent evaluation stops at a score. Potato 2.6's trajectory_edit schema lets annotators rewrite a wrong step instead of rating it, and exports each correction as supervised fine-tuning targets and DPO preference pairs. - [How to Get Reliable Labels on Agent Trajectories](https://www.potatoannotator.com/blog/annotating-agent-trajectories-reliably.md): Annotating an agent's multi-step trace is harder than labeling a tweet. A guide to designing the taxonomy, measuring step-level agreement, and adjudicating disagreement, with a Potato config. - [Can You Trust Your LLM Judge? Calibrating LLM-as-Judge Against Humans](https://www.potatoannotator.com/blog/trust-your-llm-judge-calibration.md): Using an LLM to grade model outputs is easy. Knowing whether to believe it is the hard part. A walk through Potato 2.6's blind human calibration: k-sample voting, Cohen's and Fleiss' kappa, and expected calibration error. - [Bringing Qualitative Coding to Potato: Codebooks, Memos, and In-Vivo Codes](https://www.potatoannotator.com/blog/qualitative-coding-with-potato-qda-mode.md): A look at QDA Mode, the upcoming Potato 2.6 workspace for qualitative data analysis: a living codebook, in-vivo coding, analyst memos, cases, and full-text search over a whole corpus. - [LLM Annotators vs Humans: When to Automate a Labeling Job and When Not To](https://www.potatoannotator.com/blog/llm-annotators-vs-humans.md): A practical guide to deciding when an LLM can annotate your data, where model annotators fail, and how to combine automation with human verification in Potato. - [Choosing an Open-Source Annotation Tool in 2026](https://www.potatoannotator.com/blog/choosing-an-annotation-tool-2026.md): How to pick an open-source data annotation tool, the questions that narrow the choice quickest, and where Potato fits among Label Studio, Prodigy, Doccano, brat, and Argilla. - [Codebooks for AI Annotators: Turning a Coding Scheme into a Reliable LLM Labeler](https://www.potatoannotator.com/blog/codebooks-for-ai-annotators.md): How to write an annotation codebook an LLM can actually follow, validate its labels against human coders, and keep a person in the loop, with a worked Potato config. - [Finding Hallucinations with Span Annotation](https://www.potatoannotator.com/blog/finding-hallucinations-with-span-annotation.md): Catch model hallucinations and factual errors by highlighting the exact words and labeling what's wrong, MQM-style, with span annotation in Potato. - [How Many Annotators Do You Actually Need?](https://www.potatoannotator.com/blog/how-many-annotators-do-you-need.md): Deciding annotator count and overlap for an annotation project: rules of thumb for objective and subjective tasks, the coverage-versus-overlap tradeoff, and how to set it in Potato. - [How to Evaluate RAG Systems with Human Annotation](https://www.potatoannotator.com/blog/rag-evaluation-with-human-annotation.md): A practical guide to evaluating retrieval-augmented generation: score retrieval relevance and answer faithfulness separately, and mark unsupported claims with span annotation in Potato. - [Announcing Coding Agent Annotation: Evaluate Claude Code, Aider, and SWE-Agent Traces](https://www.potatoannotator.com/blog/coding-agent-annotation-with-potato.md): Potato now supports coding agent annotation with diff rendering, terminal output display, and process reward schemas. Import traces from Claude Code, Aider, and SWE-Agent. - [How to Collect Process Reward Data for Training Better Coding Agents](https://www.potatoannotator.com/blog/process-reward-models-annotation-guide.md): Step-by-step guide to collecting per-step reward signals for PRM training using Potato. Covers first-error mode, per-step annotation, and export to training pipelines. - [Watch, Pause, and Rewind: Live Coding Agent Observation in Potato](https://www.potatoannotator.com/blog/live-coding-agent-observation.md): Tutorial for setting up live coding agent observation with Ollama, Anthropic API, or Claude Agent SDK. Includes pause, rollback, branching, and trajectory export. - [Comparing AI Agents Side by Side: Binary, Scale, and Multi-Dimension Modes](https://www.potatoannotator.com/blog/pairwise-agent-comparison-guide.md): Set up pairwise agent comparison in Potato with three modes: binary preference, continuous scale, and per-dimension multi-criteria judgment with required justification. - [Per-Step Error Localization: Using Trajectory Evaluation to Find Where Agents Fail](https://www.potatoannotator.com/blog/trajectory-evaluation-error-taxonomy.md): Use Potato's trajectory_eval schema for per-step error localization with hierarchical error taxonomies, severity scoring, and running score tracking across agent traces. - [GitHub PR-Style Code Review for AI Coding Agents](https://www.potatoannotator.com/blog/coding-agent-code-review-annotation.md): Set up GitHub PR-style code review annotation in Potato with inline diff comments, file-level quality ratings, and approve or reject verdicts for coding agent output. - [Potato vs LangSmith and Langfuse for Agent Evaluation: A Practical Comparison](https://www.potatoannotator.com/blog/potato-vs-langsmith-agent-evaluation.md): Compare Potato with LangSmith, Langfuse, Labelbox, and Scale AI for agent evaluation: trace rendering, per-step and multi-agent annotation, multimodal-agent review, coding agents, live observation, pricing, and self-hosting. - [MT-Bench-Style Rubric Evaluation for AI Agents in Potato](https://www.potatoannotator.com/blog/rubric-evaluation-mt-bench-style.md): Set up multi-criteria rubric evaluation with custom criteria, configurable rating scales, and dimension weights for systematic AI agent evaluation using Potato's rubric_eval. - [Potato 2.4.0: Web Agent Annotation, Live Evaluation, and HuggingFace Integration](https://www.potatoannotator.com/blog/potato-2-4-release.md): Potato 2.4.0 ships web agent trace review, real-time live agent evaluation, an LLM chat sidebar, HuggingFace Hub export, webhooks, SSO/OAuth, and five active learning strategies. - [Potato 2.3: Agentic Annotation, Solo Mode, and Best-Worst Scaling](https://www.potatoannotator.com/blog/potato-2-3-release.md): Potato 2.3.0 introduces agentic annotation with 12 trace format converters, Solo Mode for human-LLM collaborative labeling, Best-Worst Scaling, SSO/OAuth, Parquet export, and 15 demo projects. - [Evaluating AI Agents: Human Annotation of Agent Traces](https://www.potatoannotator.com/blog/evaluating-ai-agents-with-potato.md): How to set up human evaluation of AI agent output in Potato, from trace import through trace-level and per-step schema design to analyzing the exported results. - [Solo Mode: How One Annotator Can Label 10,000 Examples](https://www.potatoannotator.com/blog/solo-mode-tutorial.md): Step-by-step tutorial on Potato's Solo Mode: label a seed set yourself, let an LLM label the rest, and use the accuracy checks to decide when it is good enough to stop. - [Annotating Web Browsing Agents: From WebArena Traces to Human Evaluation](https://www.potatoannotator.com/blog/web-agent-annotation-guide.md): How to use Potato's web agent trace display to evaluate autonomous web browsing agents, with step-by-step screenshots, SVG overlays, and per-step annotation schemas. - [Potato 2.2: Events, Entity Linking, Export, and 55 Survey Instruments](https://www.potatoannotator.com/blog/potato-2-2-release.md): Potato 2.2.0 adds 9 new annotation schemas, a pluggable export system, MACE competence estimation, 55 validated survey instruments, and remote data sources. - [Using Visual AI to Speed Up Image and Video Annotation](https://www.potatoannotator.com/blog/visual-ai-annotation-guide.md): Set up AI-powered object detection, pre-annotation, and classification for image and video tasks with YOLO, Ollama, OpenAI, and Claude. - [Potato 2.1: Instance Display, Visual AI, and Span Linking](https://www.potatoannotator.com/blog/potato-2-1-release.md): Potato 2.1.0 brings the instance display system, visual AI support for image and video annotation, span linking, multi-field spans, and layout customization. - [Running Annotation Studies on Prolific](https://www.potatoannotator.com/blog/prolific-integration.md): Integrate Potato with Prolific for crowdsourced annotation, completion URL setup, participant ID tracking, attention checks, payment configuration, and quality filtering. - [Migrating from Label Studio to Potato](https://www.potatoannotator.com/blog/label-studio-migration.md): Migrate from Label Studio to Potato, convert project configs, annotation schemas, and exported data with this step-by-step guide covering common annotation types. - [Polygon Annotation for Segmentation Tasks](https://www.potatoannotator.com/blog/polygon-annotation-guide.md): Configure polygon drawing tools in Potato for image segmentation tasks, with tips for complex shapes, overlapping regions, multi-class polygons, and export to COCO format. - [Medical Image Annotation with Potato](https://www.potatoannotator.com/blog/medical-imaging-annotation.md): Best practices for annotating medical images in Potato, DICOM display, radiology report labeling, adverse event extraction, and IRB-compliant self-hosted deployment. - [Legal Document Annotation Best Practices](https://www.potatoannotator.com/blog/legal-document-annotation.md): Annotate legal documents in Potato, contracts, court filings, and regulatory text, with span labeling, entity extraction, and privacy-first self-hosted deployment. - [Creating a Sentiment Analysis Task](https://www.potatoannotator.com/blog/sentiment-analysis-tutorial.md): Build a complete sentiment classification task with radio buttons, tooltips, and keyboard shortcuts for efficient labeling. - [Deploying to Amazon Mechanical Turk](https://www.potatoannotator.com/blog/mturk-deployment.md): Run Potato annotation tasks on Amazon Mechanical Turk, HIT configuration, qualification tests, approval and rejection workflows, bonus payments, and quality monitoring. - [Automatic Keyword Highlighting](https://www.potatoannotator.com/blog/keyword-highlighting-setup.md): Configure AI-powered keyword highlighting in Potato to draw annotator attention to important terms. Covers OpenAI, Claude, and custom keyword list configuration. - [Introducing Potato 2.0: AI-Powered Annotation](https://www.potatoannotator.com/blog/introducing-potato-2-0.md): Potato 2.0 ships AI-powered pre-annotation with OpenAI and Claude, multimedia support for audio and video, active learning, bounding box annotation, and a redesigned UI. - [Image Comparison and Preference Tasks](https://www.potatoannotator.com/blog/image-comparison-preference.md): Build side-by-side image comparison tasks in Potato for preference ranking, A/B testing, and visual quality assessment, with randomized order and pairwise scoring. - [Exporting Annotations to Hugging Face Datasets](https://www.potatoannotator.com/blog/exporting-to-huggingface.md): Convert Potato annotations to HuggingFace Datasets format, covering JSON export structure, dataset card generation, Hub upload, and integration with transformers training. - [Drawing Bounding Boxes for Object Detection](https://www.potatoannotator.com/blog/bounding-box-annotation.md): Set up bounding box annotation for computer vision in Potato, configure label colors, minimum box size, multi-class support, validation rules, and COCO/YOLO export. - [Active Learning in Potato: Prioritizing the Items Worth Labeling](https://www.potatoannotator.com/blog/active-learning-efficiency.md): How Potato reorders the annotation queue by classifier uncertainty, with config for the sampling settings, the cold start, and keeping enough random items in the mix. - [Video Annotation with Frame-by-Frame Controls](https://www.potatoannotator.com/blog/video-frame-annotation.md): Set up video annotation in Potato with frame-by-frame navigation, timestamp markers, temporal event labeling, and per-segment classification using radio or span schemas. - [Speaker Diarization Annotation](https://www.potatoannotator.com/blog/speaker-diarization-annotation.md): Build a speaker identification task in Potato with interactive audio waveforms, timestamp markers, speaker label assignment, and inter-annotator agreement measurement. - [Pronunciation Assessment Annotation](https://www.potatoannotator.com/blog/pronunciation-assessment.md): Build a pronunciation quality annotation task in Potato with audio playback, waveform visualization, Likert rating scales, and per-recording free-text feedback fields. - [Custom HTML Templates in Potato](https://www.potatoannotator.com/blog/custom-html-templates.md): Build custom annotation interfaces in Potato using HTML templates, CSS styling, and JavaScript, side-by-side layouts, custom widgets, and embedded media displays. - [Content Moderation Annotation Setup](https://www.potatoannotator.com/blog/content-moderation-annotation.md): Configure Potato for toxicity detection, hate speech classification, and sensitive content labeling with annotator wellbeing in mind. - [Quality Control for Crowdsourced Annotation](https://www.potatoannotator.com/blog/quality-control-strategies.md): Best practices for ensuring annotation quality in annotation projects, including practical strategies you can implement with and beyond Potato. - [Measuring Inter-Annotator Agreement](https://www.potatoannotator.com/blog/inter-annotator-agreement.md): Calculate and interpret Cohen's Kappa, Fleiss' Kappa, and Krippendorff's Alpha for Potato annotation projects, with Python code examples and interpretation guidelines. - [Getting Started with Potato in 5 Minutes](https://www.potatoannotator.com/blog/getting-started-5-minutes.md): Set up your first Potato annotation project in 5 minutes, from pip install to a running server with a YAML config, sample data, and your first labeled instances. - [Audio Event Detection and Tagging](https://www.potatoannotator.com/blog/audio-event-detection.md): Set up annotation for detecting specific sounds like speech, music, applause, or environmental noises with timestamp spans. - [Deploying Potato on a Server](https://www.potatoannotator.com/blog/server-deployment-guide.md): Deploy Potato to production with Docker, nginx, SSL, and cloud platforms including AWS, GCP, and Azure. Includes systemd service configuration and scaling tips. - [Music Genre Classification Annotation](https://www.potatoannotator.com/blog/music-genre-classification.md): Create a music annotation task in Potato with waveform visualization, 30-second preview playback, hierarchical genre label trees, and pairwise preference comparisons. - [Integrating LLMs for Smart Annotation Hints](https://www.potatoannotator.com/blog/llm-integration-guide.md): Integrate OpenAI, Claude, or Gemini with Potato to provide intelligent label hints, pre-annotation suggestions, and keyword highlights that speed up annotator throughput. - [Image Classification with Potato](https://www.potatoannotator.com/blog/image-classification-tutorial.md): Set up image classification in Potato with thumbnail previews, zoom controls, multi-label checkbox schemas, and keyboard shortcuts for high-throughput visual annotation. - [Multi-Object Tracking Annotation](https://www.potatoannotator.com/blog/multi-object-tracking.md): An overview of multi-object tracking annotation concepts and how Potato's video annotation capabilities can support basic tracking workflows. - [Understanding Potato Data Formats](https://www.potatoannotator.com/blog/data-format-guide.md): How Potato reads and writes JSON and JSONL: structuring instances for text, image, audio, video, and multimodal annotation, with a working config example for each. - [Building Your First NER Annotation Task](https://www.potatoannotator.com/blog/building-ner-task.md): Step-by-step guide to building a named entity recognition task in Potato, span annotation config, label colors, overlapping spans, keyboard shortcuts, and CoNLL export. - [Classifying Emotions in Speech](https://www.potatoannotator.com/blog/audio-emotion-classification.md): Create an audio emotion classification task in Potato with interactive waveform display, playback speed controls, Likert scales, and configurable emotion label sets. - [Setting Up Audio Transcription Review](https://www.potatoannotator.com/blog/audio-transcription-task.md): Configure an audio transcription review task in Potato with waveform visualization, variable-speed playback, and inline text correction interfaces for ASR quality evaluation. - [Potato Featured at EMNLP 2022](https://www.potatoannotator.com/blog/potato-emnlp-2022.md): Our paper on Potato was accepted at EMNLP 2022. Learn about the research behind the tool and how to cite it in your work. ## Optional - [Full documentation as one file](https://www.potatoannotator.com/llms-full.txt): every doc page concatenated as Markdown. - [Source code on GitHub](https://github.com/davidjurgens/potato): the Potato repository. - [PyPI package](https://pypi.org/project/potato-annotation/): `pip install potato-annotation`.