Agreement Cannot Catch Rubber-Stamping
When annotators accept model pre-labels without reading them, inter-annotator agreement goes up. Every quality measure computed from the annotations improves. Timing is the only signal left.
News, tutorials, and updates from the Potato team.
When annotators accept model pre-labels without reading them, inter-annotator agreement goes up. Every quality measure computed from the annotations improves. Timing is the only signal left.

Potato can now record how annotators produce free-text answers, without recording what they type, and turn the pauses, revisions and pastes into auditable flags for responses that were pasted rather than written.
A full walkthrough from a folder of Whisper output to labeled speaker turns: choosing an annotation unit, handling diarization, writing the config, running the task, and exporting with the time alignment intact.



Get notified about new tutorials, releases, and community highlights.
Majority vote throws away most of what your annotators told you. Item response theory gives every label a posterior probability and a confidence interval, estimates annotator ability and item difficulty, and finds broken codebook entries, with no gold labels and no LLM.
Potato 2.7 adds labels with error bars, live norming rooms, counterfactual boundary probes, peer-prediction scoring, local voice rationales, one-command dataset reports, and mobile annotation, none of which need an LLM. Plus multi-agent and multimodal agent evaluation.

A walkthrough of human evaluation for computer-use and GUI agents in Potato: judging each action, checking click grounding on the screenshot, and reviewing tool calls one at a time.
How to find why a multi-agent LLM system failed using Potato: the interaction graph, failure attribution, handoff review, per-agent scorecards, tool-contention timeline, and emergent-behavior tagging.

Human review time is the scarcest resource in agent evaluation. Potato 2.6 pairs a signal-based triage queue with judge-human alignment so the worst traces reach people first and your LLM judge keeps getting better.
Annotator identity shapes labels on subjective tasks, so demographics are worth collecting, but only with consent and a reason. What to ask, what to leave alone, and how to run the flow in Potato.



A practical guide to deciding when an LLM can annotate your data, where model annotators fail, and how to combine automation with human verification in Potato.
How to pick an open-source data annotation tool, the questions that narrow the choice quickest, and where Potato fits among Label Studio, Prodigy, Doccano, brat, and Argilla.



Potato now supports coding agent annotation with diff rendering, terminal output display, and process reward schemas. Import traces from Claude Code, Aider, and SWE-Agent.
Step-by-step guide to collecting per-step reward signals for PRM training using Potato. Covers first-error mode, per-step annotation, and export to training pipelines.




Compare Potato with LangSmith, Langfuse, Labelbox, and Scale AI for agent evaluation: trace rendering, per-step and multi-agent annotation, multimodal-agent review, coding agents, live observation, pricing, and self-hosting.
Set up multi-criteria rubric evaluation with custom criteria, configurable rating scales, and dimension weights for systematic AI agent evaluation using Potato's rubric_eval.


How to use Potato's web agent trace display to evaluate autonomous web browsing agents, with step-by-step screenshots, SVG overlays, and per-step annotation schemas.
Potato 2.2.0 adds 9 new annotation schemas, a pluggable export system, MACE competence estimation, 55 validated survey instruments, and remote data sources.


Annotate legal documents in Potato, contracts, court filings, and regulatory text, with span labeling, entity extraction, and privacy-first self-hosted deployment.
Best practices for annotating medical images in Potato, DICOM display, radiology report labeling, adverse event extraction, and IRB-compliant self-hosted deployment.




Build side-by-side image comparison tasks in Potato for preference ranking, A/B testing, and visual quality assessment, with randomized order and pairwise scoring.
Potato 2.0 ships AI-powered pre-annotation with OpenAI and Claude, multimedia support for audio and video, active learning, bounding box annotation, and a redesigned UI.



Build custom annotation interfaces in Potato using HTML templates, CSS styling, and JavaScript, side-by-side layouts, custom widgets, and embedded media displays.
Build a pronunciation quality annotation task in Potato with audio playback, waveform visualization, Likert rating scales, and per-recording free-text feedback fields.




Calculate and interpret Cohen's Kappa, Fleiss' Kappa, and Krippendorff's Alpha for Potato annotation projects, with Python code examples and interpretation guidelines.
Best practices for ensuring annotation quality in annotation projects, including practical strategies you can implement with and beyond Potato.




An overview of multi-object tracking annotation concepts and how Potato's video annotation capabilities can support basic tracking workflows.
Create an audio emotion classification task in Potato with interactive waveform display, playback speed controls, Likert scales, and configurable emotion label sets.