Skip to content
Guides5 min read

Agreement Cannot Catch Rubber-Stamping

When annotators accept model pre-labels without reading them, agreement rises and adjudication falls. Timing records separate review from rubber-stamping.

Potato Team

When annotators accept model pre-labels without reading them, inter-annotator agreement goes up, so agreement cannot tell a rubber-stamping annotator from a careful one. The difference shows only in how the work was done, and timing is the signal that separates them.

Rubber-stamping here means accepting a model's pre-label without reading it. Consider a model that pre-labels a corpus and three annotators who review it. Two of them read each suggestion and correct what is wrong, and the third accepts everything without looking.

Compute the usual quality metrics on that study:

  • Inter-annotator agreement goes up, because the third annotator matches the model exactly, and so does most of what the other two accepted.
  • Gold-standard accuracy goes up wherever the model is right, which is most of the time.
  • Adjudication volume goes down, because there is less to adjudicate.

Every dashboard improves. On agreement, the annotator who did no work is indistinguishable from the two who did, and looks like the best annotator, because they never introduce an idiosyncratic correction. Gold items catch the rubber-stamper only where the model is wrong, and the model is right on most items.

Chance-corrected agreement under pre-labeling

Chance-corrected agreement was designed for independent judgments, and that assumption is what makes the chance baseline meaningful. Pre-labeling breaks the assumption. Two annotators who both accepted the same suggestion shared a source rather than independently arriving at the same answer. The coefficient is still computable, and it no longer means what its interpretation says it means.

Model pre-labeling is a standard feature of current annotation platforms, including Label Studio's pre-annotations, CVAT's automated annotation, and V7's Auto-Annotate. Wherever it is turned on, the independence assumption fails on every item a model labeled first, and the coefficient counts agreement with the model as agreement between people.

Timing signals for rubber-stamped pre-labels

Nothing in the annotation distinguishes rubber-stamping from careful review, because in the cases where the model was right, the two produce byte-identical output. There is no content signal to find, so the difference shows only in the process. A person cannot inspect a mask boundary and decide in 300 ms, except perhaps once, on an obvious object. As a median across many items, a 300 ms decision is not fast expertise.

Potato therefore records time per shape, stroke dynamics, revision counts, zoom behavior, and AI-suggestion accept latency, the time from a suggestion being rendered to being accepted. Accept latency is the signal that most often changes a decision.

Recording process without recording content

Recording how annotators work invites a fair objection about surveillance. Potato answers it in the design, because the event record cannot reconstruct the annotation. An event carries a timestamp, an action, a geometry kind, and one integer, and never coordinates. The absence is a structural property rather than a policy, since there is no field a coordinate could go in, and a test asserts it. Annotators also see a recording notice, and disclosure defaults to on, which draws the line between a research instrument and surveillance.

Keystroke logging applies the same principle to text. It records timing and edit structure and never records characters. It captures on beforeinput rather than keydown, because paste, drop, IME, dictation, and autofill all mutate a field without firing a keydown. A keydown-only logger is therefore blind to exactly the cases it exists to detect. The gap between characters that appeared and keys actually pressed is the strongest single signal.

Detection rules and the claims they support

No pre-trained classifier ships. There is no labeled corpus of rubber-stamping, and fitted coefficients presented as validated would be fabricated validation. Detection is instead a handful of transparent rules with visible evidence and per-project percentile calibration.

We also do not claim to detect AI-written annotations. "Surfaces responses that were pasted rather than typed" is what the data supports, and "detects AI-generated text" is not. The difference matters if anyone is going to act on a flag.

Reading timing signals as a distribution

Read timing as a distribution across items, never as a threshold on one item. Three rules hold up:

  1. Compare each annotator to the project, not to a constant. An easy corpus really is fast. Percentile calibration within your own project is the only defensible baseline.
  2. Look for runs. Twenty consecutive sub-500 ms accepts is a much stronger signal than twenty scattered across a session.
  3. Split by whether a suggestion was present. An annotator's own drawing speed is the natural control for their accept speed.

The base rate also matters. In a corpus where 80% of items are genuinely trivial, fast is usually correct, and a rule tuned to catch rubber-stamping will mostly catch competence. A flag is a prompt to look rather than a finding in itself.

Pre-labeled studies without timing records

If you are running a pre-labeled annotation study and not recording timing, you currently have no way to tell whether your annotators are reviewing or accepting. Your agreement statistics will not tell you, and they will look good. Pre-labeling genuinely helps, so the response is to measure what it changes rather than to stop using it. Telling review from rubber-stamping covers what to record and how to set it up in Potato.

Further reading