Skip to content
Free & Open Source

Why Potato?

Potato is free and runs on your own hardware. Here is how it holds up against ten other leading tools.

Why Potato?

Potato against ten other tools

We read the documentation for all of them. Most can do some of this work. What differs is how much of it sits behind a paid plan, and how much nobody sells at any price.

Included, freePaid tier onlyPartial, see notesNot offered
Annotation toolsVision & enterpriseAgent evaluation
CapabilityPotatoGPL-3.0Label StudioApache-2.0ProdigyProprietaryArgillaApache-2.0CVATMITEncordProprietaryLabelboxProprietaryLangSmithProprietaryLangfuseMITComet OpikApache-2.0GalileoProprietary
What it costsFreeFree; $99/mo hosted; agreement is Enterprise$390, or $490/seat (5 min.)FreeFree; $33/user/mo hostedNot publishedUsage-based credits$39/seat/moFree self-hostedFree self-hosted$100/mo
Self-hosted, free*Included, freeIncluded, freePaid tier onlyIncluded, freeIncluded, freePaid tier onlyPaid tier onlyPaid tier onlyIncluded, freeIncluded, freePaid tier only
Every feature, no paid tier*Included, freeNot offeredNot offeredIncluded, freeNot offeredNot offeredNot offeredNot offeredIncluded, freeIncluded, freeNot offered
One tool for corpora and agent traces*Included, freePaid tier onlyNot offeredNot offeredNot offeredNot offeredPaid tier onlyNot offeredNot offeredNot offeredNot offered
Agent trace annotationIncluded, freePaid tier onlyNot offeredNot offeredNot offeredNot offeredPaid tier onlyIncluded, freeIncluded, freeIncluded, freeIncluded, free
Multi-agent graph you can annotate*Included, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredPartial, see notesPartial, see notesPartial, see notesPartial, see notes
Human-labeled failure attribution*Included, freeNot offeredNot offeredNot offeredNot offeredNot offeredPartial, see notesNot offeredNot offeredNot offeredPartial, see notes
LLM judge measured against human labelsIncluded, freePaid tier onlyNot offeredNot offeredNot offeredNot offeredNot offeredIncluded, freeIncluded, freeNot offeredIncluded, free
Judge calibration error and bias auditIncluded, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Chance-corrected agreement (κ, α)*Included, freeNot offeredPaid tier onlyNot offeredNot offeredNot offeredNot offeredNot offeredPartial, see notesNot offeredNot offered
Annotator ability modeling (IRT)Included, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Confidence interval on every labelIncluded, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Gold standards and attention checksIncluded, freePaid tier onlyNot offeredNot offeredIncluded, freePaid tier onlyIncluded, freeNot offeredNot offeredNot offeredNot offered
Live norming sessions with an agreement meterIncluded, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Peer-prediction scoring, no gold labelsIncluded, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
SFT, DPO, and process-reward export*Included, freeNot offeredNot offeredPartial, see notesNot offeredPaid tier onlyPaid tier onlyPartial, see notesPartial, see notesNot offeredNot offered
Chance-corrected agreement over geometry*Included, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Agreement on when an event happens*Included, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
Drawing telemetry (accept latency, revisions)*Included, freeNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offeredNot offered
3D point clouds and cuboids*Included, freeNot offeredNot offeredNot offeredIncluded, freePaid tier onlyPaid tier onlyNot offeredNot offeredNot offeredNot offered
Runs with no outbound requests*Included, freePartial, see notesIncluded, freePartial, see notesIncluded, freeNot offeredPaid tier onlyNot offeredPartial, see notesPartial, see notesNot offered
Managed annotation workforce*Not offeredNot offeredNot offeredNot offeredPaid tier onlyPaid tier onlyPaid tier onlyNot offeredNot offeredNot offeredNot offered
Hosted, nothing to run*Not offeredPaid tier onlyPaid tier onlyIncluded, freeIncluded, freePaid tier onlyPaid tier onlyIncluded, freeIncluded, freeIncluded, freeIncluded, free

Self-hosted, free: With LangSmith, Galileo, Encord, and Labelbox, self-hosting is itself an enterprise contract. On-premise is an add-on you negotiate with sales rather than an edition you can download.

Every feature, no paid tier: Label Studio's free edition has no agreement metrics of any kind, and neither does the $99 hosted Starter plan. Agreement, benchmarks, and LLM pre-labeling are Enterprise features, priced on request. CVAT gives away honeypots and ground-truth jobs but keeps its quality reports for paying teams.

One tool for corpora and agent traces: Every agent-evaluation platform here is locked to traces. Ask one of them to label a plain text corpus for sentiment and it has nowhere to put the answer.

Multi-agent graph you can annotate: Four of these tools will draw you an agent interaction graph. All four are read-only: you look at it, you don't label it. In Potato the agents and the handoffs between them are the things a person annotates.

Human-labeled failure attribution: Galileo guesses at handoff failures with an LLM, and Labelbox lets you score trajectory steps. Neither one collects the human ground truth you would check that guess against.

Chance-corrected agreement (κ, α): The others report agreement as percent overlap, IoU, or a majority vote. Prodigy is the exception and ships Krippendorff's α and Gwet's AC2, but it has no free tier at all. Langfuse's Cohen's κ compares one human against one judge, not a pool of annotators against each other.

SFT, DPO, and process-reward export: Partial means SFT-style JSONL only, with no preference-pair or process-reward format.

Chance-corrected agreement over geometry: The strongest single claim in this table. CVAT's consensus engine and V7's consensus stage both compare annotators with raw IoU against a per-class threshold, and Label Studio Enterprise's metric list is exact match, numeric difference, IoU and span overlap. None of them corrects for chance, which matters because mean IoU reads about 0.95 on a corpus of large centred objects no matter who annotated it.

Agreement on when an event happens: Temporal IoU and boundary agreement, with the matching tolerance reported as a sweep rather than at one threshold. Agreement at 0.25 s and at 2 s are different claims about the same data.

Drawing telemetry (accept latency, revisions): Time per shape, stroke dynamics, revision counts and AI-suggestion accept latency. This matters because agreement goes up when annotators rubber-stamp pre-labels, so timing is the only signal that separates review from acceptance. No coordinates are recorded.

3D point clouds and cuboids: CVAT also does this, and the 3D specialists (Segments.ai, Kognic, Supervisely) go deeper, with point-cloud sequences and track propagation Potato does not have. Potato's addition is exact rotated 3D IoU between annotators.

Runs with no outbound requests: Every stylesheet, script, font and icon serves from the install across all 14 templates, verified live at 62 requests and zero external. Models are a one-time transfer. Partial means self-hostable, but the rendered page still fetches third-party assets.

Managed annotation workforce: We don't have one. You bring your own annotators or recruit through Prolific. If you need 50,000 items labeled by next month and nobody to label them, Labelbox, Encord, and Scale sell exactly that, and we don't.

Hosted, nothing to run: Potato is self-hosted only. For a lab with a spare server that's the whole appeal. If you don't want to run anything at all, it's a real cost and you should weigh it.

Checked against public documentation in August 2026. These products change quickly, so if we have a cell wrong we'd like to know. Tell us what we got wrong.

Seven reasons teams switch

Free, with no tiers

There is no paid plan, because there is no plan. Every feature on this page behaves the same whether you're one graduate student or a company of four hundred. Nothing is metered and nobody asks for a card.

It measures the labeling, not just the labels

Most tools store what an annotator clicked. Potato also estimates how reliable each annotator is and how hard each item was, so a label comes back with a confidence interval instead of a majority vote pretending to be a fact.

Open source

GPL-3.0, all of it on GitHub. You can read how the statistics are computed, change the parts you disagree with, and keep running your fork long after we lose interest. Several tools on this page are open source too. Fewer of them give you the whole product.

Your data stays put

Potato runs on your laptop or your server and never phones home. That is what makes it usable for medical notes and unpublished corpora that legal won't let near someone else's cloud.

Built in a research lab

Potato came out of the University of Michigan, with papers at ACL 2026 and EMNLP 2022 and a best demo award at HCOMP 2024. If a reviewer asks how your agreement numbers were computed, the method is published and you can cite it.

It's a YAML file

A working annotation task is about twenty lines of config. There is no plugin API to learn and no build step. Whoever writes your codebook can set up the study without waiting on an engineer.

Agent evaluation

The broadest agent annotation support of any tool we found: 15 trace formats, multi-agent runs drawn as an interaction graph a person can actually label, schemas for computer-use and voice and video, coding-agent diffs, process-reward annotation, and live observation with rollback. Other platforms show you an agent graph. Potato is where somebody labels it.

Who uses Potato

Academic researchers

Free, citable, and self-hosted, which is usually what an IRB and a grant budget both need.

Industry teams

Runs on your own infrastructure. No per-seat licence, and no vendor sitting on your labels.

Startups and small teams

Start this afternoon with a pip install instead of a procurement cycle.

Privacy-sensitive projects

Works fully offline, which is why it ends up on medical notes and student records that can't go to a vendor's cloud.

Educators and students

Free for a whole class. No seats to buy, no licences to chase at the start of term.

Individual developers

Build a training set for your own model without signing up for anything.

Try it

One pip install and you're annotating. There is nothing to sign up for.