Skip to content
Guides5 min read

Shipping Segmentation That Runs in a Browser Tab

Click-to-segment and text prompting run in the browser with no GPU, using a verified encoder contract, a measured quantization, and a hand-written tokenizer.

Potato Team

Potato runs MobileSAM for click-to-segment and Grounding DINO for open-vocabulary text prompting in the annotator's browser through ONNX Runtime Web, with no GPU. Getting there took an encoder contract verified against ground truth, a quantization chosen by measurement, and a tokenizer written by hand.

pip install potato-annotation should give you working segmentation with no GPU, no new Python dependency, and no outbound network at annotation time. The model weights are a one-time download on top of that install. potato download-models onnxruntime fetches 13.5 MB and potato download-models mobile_sam fetches 45 MB, and an air-gapped site runs both on a connected machine and copies potato/models/ across. The offline constraint matters because several research groups deploy Potato air-gapped, where a model fetched at click time is a missing feature rather than a slow one. Both of Potato's image-assist models therefore run client-side. The first three sections below cover decisions from that work that generalize past this project.

Plausible readings of the SAM encoder contract

SAM's image encoder takes a preprocessed tensor, and the specification leaves room for interpretation in normalization, resize convention, padding, and channel order.

We implemented one reading and got masks that looked right. Centroid error against ground truth was 70 to 148 pixels, the sort of error that reads as a slightly sloppy annotator rather than as a broken pipeline. Three separate plausible readings all produced confident, plausible, wrong masks. The correct reading lands at 0.1 px.

The lesson is that a wrong reading here is invisible without ground truth, rather than that the spec needs a more careful reading. The masks are the right shape, in roughly the right place, and every downstream check passes. Only a numeric comparison against known truth separates the four candidates, and only one of them is right.

The contract is therefore pinned by a test against real weights rather than documented in a comment. A comment would have been just as true and would not have caught the regression.

Quantization chosen by measurement

A 686 MB model has to be quantized to run in a browser, and the choice of quantization is usually whatever the export tool defaults to. We measured two exports against the full-precision Grounding DINO export:

ExportBox IoU vs full precisionSize
q4f160.972151 MB
int80.874201 MB

q4f16 holds more of the original geometry (0.972 box IoU against 0.874) and is 50 MB smaller. int8 is the conventional choice and would have shipped a worse model in a larger file. The measurement took an afternoon, and without it the choice would have been permanent, because nobody re-examines a quantization choice once boxes are appearing on screen.

We verified the result live afterwards on the COCO two-cat photo. Both cats were detected, at 0.724 and 0.688, with boxes matching the Python reference implementation to four decimal places.

Writing the tokenizer by hand

Grounding DINO needs its caption tokenized the way BERT does it. The obvious move is to pull in transformers.js, but that library is about 2 MB of JavaScript for one function, on a page that already loads a 151 MB model. We implemented WordPiece directly in about 200 lines and checked it token for token against Hugging Face's tokenizers.

The saved bytes matter less than correctness. A tokenizer that disagrees with the model's training tokenizer produces subtly wrong detections that look like a threshold problem, and you would spend a day tuning box_threshold before suspecting the tokenizer. The token-level equivalence test compares the JavaScript tokenizer with Hugging Face's on a fixed set of prompts that exercise subword splits, the period separator, accents, hyphens, digits, casing, and characters no vocabulary contains. Any divergence on those inputs fails the test before it can surface as a mislabeled box, and a caption that breaks in a new way belongs in that set.

Server-side video mask propagation

Video mask propagation runs server-side, as a deliberate exception rather than an inconsistency, because the costs differ. Click-to-segment costs one cheap decoder pass per click, against an encoding computed once per image and cached. Propagation costs a full model pass per frame, the model is 181 MB across five graphs, and the video is already on the server. A hundred frames in a browser tab is minutes of a frozen page.

We keep a lighter in-browser carry-forward path that re-prompts frame to frame, and we are careful not to describe the two as the same capability. "SAM 2 propagation runs in your browser" would be a better sentence and a false one.

Recommendations for running models in the browser

If you are building something similar, we would recommend four things:

  1. Get ground truth before you get output. Plausible wrong output is the failure mode, and it survives review.
  2. Measure the quantization. The default is a choice someone else made for a different constraint.
  3. Test the preprocessing against the reference implementation numerically. Visual inspection cannot distinguish 0.1 px from 148 px of error at a glance.
  4. Separate the encoder from the decoder and cache the expensive half. The separation is the entire reason clicking feels interactive.

Further reading