Potato 2.8: Computer Vision, 3D, Robot Episodes, and Spatial Agreement
Potato 2.8 adds computer vision, deep zoom, 3D point clouds, depth maps, robot episodes, and video-model evaluation, with chance-corrected agreement for each.
Potato 2.8 covers images, video, gigapixel scans, 3D point clouds, depth maps, robot episodes, generative-video rollouts, and vision-language grounding, and for every one of them it reports whether the annotators agreed. Before this release, Potato was a text annotation platform with an image mode.
Drawing tools are common. CVAT, Label Studio, Roboflow, V7, and Supervisely all draw boxes and polygons, several of them better than Potato does. A chance-corrected reliability statistic over spatial labels is scarce, and we have not found another annotation platform that reports one. CVAT's consensus engine and V7's consensus stage both compare annotators by raw intersection over union (IoU) against a per-class threshold, with no chance correction.
Potato 2.8
Upgrade notes
Wheels up to 2.7.1 declared their package data with a single-level glob, so 27 templates stored in subdirectories never made it into the published package. Solo mode's setup routes, the admin dashboard, judge calibration, and the corpus map were all broken for anyone who ran pip install. Anyone running from a git checkout never saw the problem, which is also why the test suite never caught it. @0x6362 reported and fixed it in PR #164. If you installed from PyPI, upgrade before anything else:
pip install --upgrade potato-annotationImage annotation keyboard shortcuts now follow V7 conventions by default (b brush, r rectangle, f fill, k keypoint, v select). Set keybinding_profile: legacy to restore the old bindings for a study already in the field.
Chance-corrected agreement for spatial labels
Take a corpus where every image holds one large, centered object. Two annotators who both draw a box in the middle will overlap heavily, so mean IoU comes out around 0.95, and so will an annotator who drew a box in the middle without looking at the image. The number measures how easy the task is rather than how much the annotators agree, which is the problem chance correction was invented to solve for categorical labels. Potato reports σ, which is Krippendorff's α in its own 1 − D_o/D_e form, generalized to an arbitrary distance, with the chance baseline estimated empirically by comparing annotations of different items. σ = 0 means annotators agree no more than they would on unrelated images.
Our original plan was Krippendorff's α over 1 − IoU, and it does not work because IoU distance saturates. Two randomly paired boxes almost always have IoU 0, so expected disagreement sits at about 1 and α degenerates into 1 − mean distance with no chance correction left in it. Braylan, Alonso, and Lease (WWW 2022) found the same thing empirically. They report that for bounding boxes α ranks plain L2 above IoU and GIoU, which inverts the ordering practitioners give.
Potato also reports spatial agreement as three numbers rather than one, covering whether annotators found the same objects, called them the same thing, and put them in the same place. An annotator who finds everything and mislabels it has a different problem from one who labels correctly and boxes loosely, and the two problems have different fixes.
Adjudication bug in image agreement
Adjudication compared the keys of an annotation, and an image schema stores everything under one key named _data. Every image pair therefore scored 1.0 agreement, so two annotators who agreed on nothing looked unanimous, and no image was ever routed for review. The bug is fixed, and it is the reason the agreement layer is built into each annotation type rather than added on top.
In-browser segmentation and text prompting
Clicking an object returns a mask in about 130 ms, with no GPU and no network call per click, using MobileSAM under ONNX Runtime Web. Typing a phrase boxes every match, using Grounding DINO. Grounding DINO is released under Apache-2.0, so text prompting ships in every install with no license to accept. SAM 3 also labels from text, but it is under Meta's SAM License, which permits commercial use and adds acceptable-use restrictions. Potato bundles no SAM 3 weights and offers it only as a server endpoint that will not download without --accept-licence.
Two details from building these features apply beyond Potato.
The encoder contract was checked against real weights. SAM's input specification admits several plausible readings, and three of them produce confident, plausible masks that are wrong, with 70 to 148 pixels of centroid error. An error of that size looks like a slightly sloppy annotator rather than a broken pipeline. The correct reading lands at 0.1 px, and only a check against real weights and real ground truth separates them.
The quantization was chosen by measurement. Against the 686 MB full-precision Grounding DINO export, q4f16 holds box IoU 0.972 where int8 holds 0.874, and q4f16 is also 50 MB smaller. Taking the conventional choice would have shipped a worse model in a larger file.
SAM 2 video tracking and occluded frames
When you draw a mask on one frame and press Track forward, SAM 2 follows the object through the frames that follow. Against known ground truth, per-frame IoU was 0.974 to 0.979 with no decay from the first frame to the last, at about 1.32 s per frame on CPU. SAM 2 tracking runs server-side by design, because its cost is per frame rather than per prompt, and a hundred frames of a video model in a browser tab would freeze the page for minutes. Potato also has a lighter in-browser carry-forward path, and the two are separate features.
The most useful property in an annotation loop is that the model decides for itself when the object is occluded and returns an empty frame rather than guessing. A plausible wrong mask on a hidden frame is work to undo. An empty frame is immediately readable, and it is also the correct answer.
Point clouds, depth maps, and robot episodes
spatial_annotation reads PCD, PLY, LAS, KITTI .bin, and .xyz, with octree level of detail, orthographic slab panels, and calibration that projects every 3D box into every camera image so annotators can verify in 2D while editing in 3D.
Rotation is stored as a quaternion rather than a yaw angle, which makes KITTI import lossless. The quaternion keeps the camera-to-lidar mounting tilt of about 0.85° that a yaw-only field discards silently. Export back to KITTI reports how much orientation it had to drop rather than flattening the box quietly.
Depth maps load from 16-bit PNG and TIFF, NPY, PFM, and EXR, with a readout in meters under the cursor. Potato paints zero magenta rather than rendering it as "very close," because zero is the near-universal no-return code, and a stereo rig returns it wherever it finds no texture to match, such as a blank wall.
episode_annotation puts N synchronized video streams and M robot time-series lanes on one timeline. It records three outcomes rather than two, because partial success is the modal result in real robot data, and forcing it into a binary destroys the signal that makes the dataset worth annotating. Lane downsampling preserves minima and maxima, so a one-frame force spike survives the trip to a 300-pixel lane. The spike matters because a missed grasp is visible in the force trace several frames before it is obvious on camera.
Break-point evaluation for generated video
rollout_evaluation shows 2 to N generated videos on one clock and asks the annotator to mark the frame at which the world stops making sense, then tag which physical or causal property broke. A rating of 3 out of 5 for "physical plausibility" cannot be checked, cannot be localized, and cannot be used to fix anything. A frame index plus a category can be all three, and two annotators' answers are two points on a line, so a real agreement statistic applies. Potato reports break-point agreement as detection, localization, and category, with the matching tolerance shown as a sweep rather than one number, because agreement at 0.25 s and agreement at 2 s are different claims about the same data.
Drawing telemetry for pre-labeled annotations
Pre-labeling makes annotators faster, and it also makes rubber-stamping frictionless. A suggestion appears, the annotator clicks accept, and the dataset becomes a record of a model agreeing with itself. Nothing in the annotation distinguishes rubber-stamping from careful review. The geometry is identical, and every quality measure looks better, inter-annotator agreement included, because annotators who all accept the same pre-label are unanimous by construction.
The difference shows up in how each shape was made. Drawing telemetry records time per shape, stroke dynamics, revision counts, and AI-suggestion accept latency. It never records coordinates, as a structural property of the event record rather than a policy. As one-click auto-labeling becomes the norm, that record of how a shape was made is what lets you tell review from rubber-stamping, because the finished geometry looks the same either way.
Also in 2.8
The release also includes these changes:
- Threaded conversations. The
dialoguedisplay renders reply structure from each turn'sreply_to, andpotato convokitimports any ConvoKit corpus and exports back, with noconvokitdependency. - 15 import and 29 export formats, 11 of them round-tripping. Darwin works in both directions, which matters if you are leaving a platform rather than joining one.
- Live database ingestion. Rows created after startup become annotatable within
poll_interval_secondswith no restart. - Machine-readable specs. A JSON Schema covers 159 config keys, 61 annotation types, and 24 display types, and an OpenAPI document covers 419 paths. Both are generated from the code and checked in CI.
- No outbound requests. Every asset serves from the install across all 14 templates.
styles.csshad been opening with an@importof Google Fonts, which sent every annotator's IP address to a third party on every page load, from a tool people self-host precisely so their data does not leave their infrastructure. The import had survived three previous air-gap audits, because every guard read<script src>and<link href>, and an@importinside a stylesheet is neither. - The admin Instances tab went from quadratic to linear time. With 2,000 items, it dropped from 15.1 s to 23 ms.
Where to start
Everything above is in the free, self-hosted product, and there is no paid tier. These pages cover the new features:
- Vision and spatial annotation
- Measurement and integrity
- Measuring agreement on bounding boxes
- Full release notes
References
Alexander Braylan, Omar Alonso, and Matthew Lease (2022). Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation Tasks. Proceedings of the ACM Web Conference 2022. https://doi.org/10.1145/3485447.3512242