Skip to content
Free & Open Source

Why Potato?

Potato is free and runs on your own hardware. Here is how it holds up against eleven other leading tools.

Why Potato?

Potato against eleven other tools

We read the documentation for all of them. Most can do some of this work. What differs is how much of it sits behind a paid plan, and how much nobody sells at any price.

Included, freePaid tier onlyPartial, see notesNot offered
Annotation toolsVision & enterpriseAgent evaluation
CapabilityPotatoGPL-3.0Label StudioApache-2.0ProdigyProprietaryArgillaApache-2.0CVATMITEncordProprietaryLabelboxProprietaryLangSmithProprietaryLangfuseMITComet OpikApache-2.0GalileoProprietary
What it costsFree$99/user/mo for agreement metrics$390, or $490/seat (5 min.)FreeFree; $33/user/mo for QCNot publishedUsage-based credits$39/seat/moFree self-hostedFree self-hosted$100/mo
Self-hosted, free*Included, freeIncluded, freePaid tier onlyIncluded, freeIncluded, freePaid tier onlyNot offeredPaid tier onlyIncluded, freeIncluded, freePaid tier only
Every feature, no paid tier*Included, freeNot offeredNot offeredIncluded, freeNot offeredNot offeredNot offeredNot offeredIncluded, freeIncluded, freeNot offered
One tool for corpora and agent traces*Included, freePaid tier onlyNot offeredNot offeredNot offeredNot offeredPaid tier onlyNot offeredNot offeredNot offeredNot offered
Agent trace annotationIncluded, freePaid tier onlyNot offeredNot offeredNot offeredNot offeredPaid tier onlyIncluded, freeIncluded, freeIncluded, freeIncluded, free
Multi-agent graph you can annotate*Included, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredPartial, see notesPartial, see notesPartial, see notesPartial, see notes
Human-labeled failure attribution*Included, freeNot offeredNot offeredNot offeredNot offeredNot offeredPartial, see notesNot offeredNot offeredNot offeredPartial, see notes
LLM judge measured against human labelsIncluded, freePaid tier onlyNot offeredNot offeredNot offeredNot offeredNot offeredIncluded, freeIncluded, freeNot offeredIncluded, free
Judge calibration error and bias auditIncluded, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Chance-corrected agreement (κ, α)*Included, freeNot offeredPaid tier onlyNot offeredNot offeredNot offeredNot offeredNot offeredPartial, see notesNot offeredNot offered
Annotator ability modeling (IRT)Included, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Confidence interval on every labelIncluded, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Gold standards and attention checksIncluded, freePaid tier onlyNot offeredNot offeredIncluded, freePaid tier onlyIncluded, freeNot offeredNot offeredNot offeredNot offered
Live norming sessions with an agreement meterIncluded, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Peer-prediction scoring, no gold labelsIncluded, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
SFT, DPO, and process-reward export*Included, freeNot offeredNot offeredPartial, see notesNot offeredPaid tier onlyPaid tier onlyPartial, see notesPartial, see notesNot offeredNot offered
Peer-reviewed paper describing the toolIncluded, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Managed annotation workforce*Not offeredNot offeredNot offeredNot offeredPaid tier onlyPaid tier onlyPaid tier onlyNot offeredNot offeredNot offeredNot offered
Hosted, nothing to run*Not offeredPaid tier onlyPaid tier onlyIncluded, freeIncluded, freePaid tier onlyPaid tier onlyIncluded, freeIncluded, freeIncluded, freeIncluded, free

Self-hosted, free: With LangSmith, Galileo, and Encord, self-hosting is itself an enterprise contract. Labelbox documents no self-hosted edition at all.

Every feature, no paid tier: Label Studio's free edition has no agreement metrics of any kind. Ground truth, dashboards, and LLM pre-labeling start at $99 per user per month. CVAT lists its quality-control interface as paid, and Supervisely puts consensus behind a €199 per month plan.

One tool for corpora and agent traces: Every agent-evaluation platform here is locked to traces. Ask one of them to label a plain text corpus for sentiment and it has nowhere to put the answer.

Multi-agent graph you can annotate: Four of these tools will draw you an agent interaction graph. All four are read-only: you look at it, you don't label it. In Potato the agents and the handoffs between them are the things a person annotates.

Human-labeled failure attribution: Galileo guesses at handoff failures with an LLM, and Labelbox lets you score trajectory steps. Neither one collects the human ground truth you would check that guess against.

Chance-corrected agreement (κ, α): The others report agreement as percent overlap, IoU, or a majority vote. Prodigy is the exception and ships Krippendorff's α and Gwet's AC2, but it has no free tier at all. Langfuse's Cohen's κ compares one human against one judge, not a pool of annotators against each other.

SFT, DPO, and process-reward export: Partial means SFT-style JSONL only, with no preference-pair or process-reward format.

Managed annotation workforce: We don't have one. You bring your own annotators or recruit through Prolific. If you need 50,000 items labeled by next month and nobody to label them, Labelbox, Encord, and Scale sell exactly that, and we don't.

Hosted, nothing to run: Potato is self-hosted only. For a lab with a spare server that's the whole appeal. If you don't want to run anything at all, it's a real cost and you should weigh it.

Checked against public documentation in July 2026. These products change quickly, so if we have a cell wrong we'd like to know. Tell us what we got wrong.

Seven reasons teams switch

Free, with no tiers

There is no paid plan, because there is no plan. Every feature on this page behaves the same whether you're one graduate student or a company of four hundred. Nothing is metered and nobody asks for a card.

It measures the labeling, not just the labels

Most tools store what an annotator clicked. Potato also estimates how reliable each annotator is and how hard each item was, so a label comes back with a confidence interval instead of a majority vote pretending to be a fact.

Open source

GPL-3.0, all of it on GitHub. You can read how the statistics are computed, change the parts you disagree with, and keep running your fork long after we lose interest. Several tools on this page are open source too. Fewer of them give you the whole product.

Your data stays put

Potato runs on your laptop or your server and never phones home. That is what makes it usable for medical notes and unpublished corpora that legal won't let near someone else's cloud.

Built in a research lab

Potato came out of the University of Michigan, with papers at ACL 2026 and EMNLP 2022 and a best demo award at HCOMP 2024. None of the eleven tools we compared it against has a peer-reviewed paper describing it. If you have to cite your instrument in a methods section, that matters.

It's a YAML file

A working annotation task is about twenty lines of config. There is no plugin API to learn and no build step. Whoever writes your codebook can set up the study without waiting on an engineer.

Agent evaluation

The broadest agent annotation support of any tool we found: 13 trace formats, multi-agent runs drawn as an interaction graph a person can actually label, schemas for computer-use and voice and video, coding-agent diffs, process-reward annotation, and live observation with rollback. Other platforms show you an agent graph. Potato is where somebody labels it.

What Teams Say

We evaluated Labelbox and Scale AI, but Potato let us get started in an afternoon without any budget approval.

ML Researcher

University Research Lab

The YAML configuration is so much simpler than wrestling with complex annotation tool UIs. Our whole team was onboarded in a day.

Data Scientist

NLP Startup

Being able to self-host was critical for our IRB-approved study with sensitive participant data.

PhD Student

Social Science Department

Who uses Potato

Academic researchers

Free, citable, and self-hosted, which is usually what an IRB and a grant budget both need.

Industry teams

Runs on your own infrastructure. No per-seat licence, and no vendor sitting on your labels.

Startups and small teams

Start this afternoon with a pip install instead of a procurement cycle.

Privacy-sensitive projects

Works fully offline, which is why it ends up on medical notes and student records that can't go to a vendor's cloud.

Educators and students

Free for a whole class. No seats to buy, no licences to chase at the start of term.

Individual developers

Build a training set for your own model without signing up for anything.

Try it

One pip install and you're annotating. There is nothing to sign up for.