Partial Is a Real Outcome: Annotating Robot Episodes
Many robot episodes end in partial success, and binary labels discard it. Annotation needs a timeline, min/max downsampling, and relabeling.
Annotate robot episodes on a timeline with three outcomes (success, partial success, and failure), because many real robot episodes end in partial success and a binary label throws that outcome away. Put the numeric signals beside the video, downsampled with min/max rather than the mean, so a one-frame force spike survives.
A robot demonstration, or episode, is several synchronized camera streams plus several numeric time series such as joint positions, gripper state, force-torque, and reward, all indexed by frame. Every question worth asking about an episode is temporal, such as when the grasp started, when it failed, and how close to done it was at each moment. The annotation interface therefore has to be a timeline rather than a form, and the sections below cover what that required in Potato.
Outcome labels with partial success
The first design decision was the outcome vocabulary. Success or failure is the natural binary, and it destroys the signal, because many real robot episodes end in partial success. The arm got the object off the table and dropped it, or placed the block in the right bin but the wrong compartment, or completed the task two seconds after the timeout.
Forcing those episodes into "failure" makes them indistinguishable from an arm that never moved. Forcing them into "success" is worse. A policy trained on either mapping learns something untrue about what happened.
Three outcomes with a failure-cause taxonomy attached make the failures analyzable instead of merely counted. The taxonomy covers missed grasp, object slipped, collision, wrong object, wrong placement, and timeout.
Signal lanes beside the video
Showing the numeric signals beside the video has published support. ATLAS (TU Wien, 2026) reports that adding proprioceptive time series to the video cut boundary error fivefold against vision-only annotation tools on a contact-rich assembly task, and that finding is why Potato renders those lanes at all.
The finding matches what you see immediately on real episodes. A missed grasp is visible in the gripper and wrist-force traces several frames before it becomes obvious on camera. The gripper closes, the force never rises, and only later does the video show an empty hand. If the lanes are behind a tab, annotators do not open them.
Min/max downsampling for signal lanes
A 30-second episode at 30 Hz is 900 samples per series, drawn into a lane about 300 pixels wide, so something has to be dropped. Mean downsampling drops exactly the wrong thing, because a one-frame force spike is the event, and averaging it with its neighbors removes it entirely. The lane looks smooth and the moment of contact is gone. Min/max-preserving downsampling keeps both extremes in each bucket, so a single-frame spike survives to the pixel. The choice is a small implementation detail, and it determines whether the lane is useful at all.
Annotation layers and their cost
Potato's three annotation layers for robot episodes are not comparable in cost:
| Layer | Cost per episode |
|---|---|
| Phase segmentation | Seconds |
| Outcome plus cause | Seconds |
| Dense progress reward | Minutes |
The layers are therefore individually configurable, and most projects enable a subset. Shipping all three as one non-negotiable form would have made the reward curve the reason nobody finished a study.
Hindsight relabeling of failed episodes
An episode that failed at the task it was recorded for is frequently a successful demonstration of some other task. If the arm was meant to place the block in the left bin and put it in the right one, the episode is a failure of the recorded instruction and a clean success of a different one.
Collecting what the episode actually demonstrates costs one text field at annotation time and is impossible to recover later. Ask for it even if you are not sure you will use it.
Agreement measures for each layer
The layers produce three different kinds of answer, so they need three measures:
- Phase boundaries: temporal IoU
- Outcome labels: Krippendorff's α
- Reward curves: ICC plus Pearson
Each measure is reported with coverage stated alongside. A correlation computed over 5% of a timeline says nothing about the other 95%, and a coefficient without its coverage will be read as though it covered everything. For boundaries, the tolerance also matters. Agreement at 0.25 s and agreement at 2 s are different claims, so the report sweeps the tolerance rather than picking one.
Import and export formats
LeRobot v2, RLDS/TFDS, HDF5 (RoboMimic and ALOHA), and ROS bags all import through the same field. Export is LeRobot v2 plus a per-frame JSONL sidecar, so a read-only public dataset can carry your annotations without being republished.
Other tools for robot and behavioral annotation
ATLAS is built for precise action boundaries in long-horizon robot episodes recorded as ROS bags or RLDS. ELAN and BORIS are established tools for behavioral coding. Rerun logs and visualizes multi-rate, multimodal recordings, and Foxglove collects, searches, and debugs robot fleet data.
Potato adds multi-annotator, web-deployed annotation with agreement statistics over the result. If one person is annotating, use ATLAS.
Further reading
References
Sergej Stanovcic, Daniel Sliwowski & Dongheui Lee (2026). ATLAS: An Annotation Tool for Long-horizon Robotic Action Segmentation. arXiv preprint arXiv:2604.26637. https://arxiv.org/abs/2604.26637