Why Potato?
Potato is free and runs on your own hardware. Here is how it holds up against eleven other leading tools.
Potato against eleven other tools
We read the documentation for all of them. Most can do some of this work. What differs is how much of it sits behind a paid plan, and how much nobody sells at any price.
| Annotation tools | Vision & enterprise | Agent evaluation | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Capability | PotatoGPL-3.0 | Label StudioApache-2.0 | ProdigyProprietary | ArgillaApache-2.0 | CVATMIT | EncordProprietary | LabelboxProprietary | LangSmithProprietary | LangfuseMIT | Comet OpikApache-2.0 | GalileoProprietary |
| What it costs | Free | $99/user/mo for agreement metrics | $390, or $490/seat (5 min.) | Free | Free; $33/user/mo for QC | Not published | Usage-based credits | $39/seat/mo | Free self-hosted | Free self-hosted | $100/mo |
| Self-hosted, free* | Included, free | Included, free | Paid tier only | Included, free | Included, free | Paid tier only | Not offered | Paid tier only | Included, free | Included, free | Paid tier only |
| Every feature, no paid tier* | Included, free | Not offered | Not offered | Included, free | Not offered | Not offered | Not offered | Not offered | Included, free | Included, free | Not offered |
| One tool for corpora and agent traces* | Included, free | Paid tier only | Not offered | Not offered | Not offered | Not offered | Paid tier only | Not offered | Not offered | Not offered | Not offered |
| Agent trace annotation | Included, free | Paid tier only | Not offered | Not offered | Not offered | Not offered | Paid tier only | Included, free | Included, free | Included, free | Included, free |
| Multi-agent graph you can annotate* | Included, free | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Partial, see notes | Partial, see notes | Partial, see notes | Partial, see notes |
| Human-labeled failure attribution* | Included, free | Not offered | Not offered | Not offered | Not offered | Not offered | Partial, see notes | Not offered | Not offered | Not offered | Partial, see notes |
| LLM judge measured against human labels | Included, free | Paid tier only | Not offered | Not offered | Not offered | Not offered | Not offered | Included, free | Included, free | Not offered | Included, free |
| Judge calibration error and bias audit | Included, free | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered |
| Chance-corrected agreement (κ, α)* | Included, free | Not offered | Paid tier only | Not offered | Not offered | Not offered | Not offered | Not offered | Partial, see notes | Not offered | Not offered |
| Annotator ability modeling (IRT) | Included, free | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered |
| Confidence interval on every label | Included, free | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered |
| Gold standards and attention checks | Included, free | Paid tier only | Not offered | Not offered | Included, free | Paid tier only | Included, free | Not offered | Not offered | Not offered | Not offered |
| Live norming sessions with an agreement meter | Included, free | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered |
| Peer-prediction scoring, no gold labels | Included, free | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered |
| SFT, DPO, and process-reward export* | Included, free | Not offered | Not offered | Partial, see notes | Not offered | Paid tier only | Paid tier only | Partial, see notes | Partial, see notes | Not offered | Not offered |
| Peer-reviewed paper describing the tool | Included, free | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered | Not offered |
| Managed annotation workforce* | Not offered | Not offered | Not offered | Not offered | Paid tier only | Paid tier only | Paid tier only | Not offered | Not offered | Not offered | Not offered |
| Hosted, nothing to run* | Not offered | Paid tier only | Paid tier only | Included, free | Included, free | Paid tier only | Paid tier only | Included, free | Included, free | Included, free | Included, free |
Self-hosted, free: With LangSmith, Galileo, and Encord, self-hosting is itself an enterprise contract. Labelbox documents no self-hosted edition at all.
Every feature, no paid tier: Label Studio's free edition has no agreement metrics of any kind. Ground truth, dashboards, and LLM pre-labeling start at $99 per user per month. CVAT lists its quality-control interface as paid, and Supervisely puts consensus behind a €199 per month plan.
One tool for corpora and agent traces: Every agent-evaluation platform here is locked to traces. Ask one of them to label a plain text corpus for sentiment and it has nowhere to put the answer.
Multi-agent graph you can annotate: Four of these tools will draw you an agent interaction graph. All four are read-only: you look at it, you don't label it. In Potato the agents and the handoffs between them are the things a person annotates.
Human-labeled failure attribution: Galileo guesses at handoff failures with an LLM, and Labelbox lets you score trajectory steps. Neither one collects the human ground truth you would check that guess against.
Chance-corrected agreement (κ, α): The others report agreement as percent overlap, IoU, or a majority vote. Prodigy is the exception and ships Krippendorff's α and Gwet's AC2, but it has no free tier at all. Langfuse's Cohen's κ compares one human against one judge, not a pool of annotators against each other.
SFT, DPO, and process-reward export: Partial means SFT-style JSONL only, with no preference-pair or process-reward format.
Managed annotation workforce: We don't have one. You bring your own annotators or recruit through Prolific. If you need 50,000 items labeled by next month and nobody to label them, Labelbox, Encord, and Scale sell exactly that, and we don't.
Hosted, nothing to run: Potato is self-hosted only. For a lab with a spare server that's the whole appeal. If you don't want to run anything at all, it's a real cost and you should weigh it.
Checked against public documentation in July 2026. These products change quickly, so if we have a cell wrong we'd like to know. Tell us what we got wrong.
Seven reasons teams switch
Free, with no tiers
There is no paid plan, because there is no plan. Every feature on this page behaves the same whether you're one graduate student or a company of four hundred. Nothing is metered and nobody asks for a card.
It measures the labeling, not just the labels
Most tools store what an annotator clicked. Potato also estimates how reliable each annotator is and how hard each item was, so a label comes back with a confidence interval instead of a majority vote pretending to be a fact.
Open source
GPL-3.0, all of it on GitHub. You can read how the statistics are computed, change the parts you disagree with, and keep running your fork long after we lose interest. Several tools on this page are open source too. Fewer of them give you the whole product.
Your data stays put
Potato runs on your laptop or your server and never phones home. That is what makes it usable for medical notes and unpublished corpora that legal won't let near someone else's cloud.
Built in a research lab
Potato came out of the University of Michigan, with papers at ACL 2026 and EMNLP 2022 and a best demo award at HCOMP 2024. None of the eleven tools we compared it against has a peer-reviewed paper describing it. If you have to cite your instrument in a methods section, that matters.
It's a YAML file
A working annotation task is about twenty lines of config. There is no plugin API to learn and no build step. Whoever writes your codebook can set up the study without waiting on an engineer.
Agent evaluation
The broadest agent annotation support of any tool we found: 13 trace formats, multi-agent runs drawn as an interaction graph a person can actually label, schemas for computer-use and voice and video, coding-agent diffs, process-reward annotation, and live observation with rollback. Other platforms show you an agent graph. Potato is where somebody labels it.
What Teams Say
“We evaluated Labelbox and Scale AI, but Potato let us get started in an afternoon without any budget approval.”
ML Researcher
University Research Lab
“The YAML configuration is so much simpler than wrestling with complex annotation tool UIs. Our whole team was onboarded in a day.”
Data Scientist
NLP Startup
“Being able to self-host was critical for our IRB-approved study with sensitive participant data.”
PhD Student
Social Science Department
Who uses Potato
Academic researchers
Free, citable, and self-hosted, which is usually what an IRB and a grant budget both need.
Industry teams
Runs on your own infrastructure. No per-seat licence, and no vendor sitting on your labels.
Startups and small teams
Start this afternoon with a pip install instead of a procurement cycle.
Privacy-sensitive projects
Works fully offline, which is why it ends up on medical notes and student records that can't go to a vendor's cloud.
Educators and students
Free for a whole class. No seats to buy, no licences to chase at the start of term.
Individual developers
Build a training set for your own model without signing up for anything.
Try it
One pip install and you're annotating. There is nothing to sign up for.