Documenting Your Annotation Dataset: Data Statements, Datasheets, and a Release Checklist
Record the curation rationale, annotator pool, guidelines, and intended use when you release an annotation dataset, and see how much of it Potato captures.
A data statement records a dataset's curation rationale, the language and its speakers, the annotators and their demographics, the guidelines, and the intended use, so downstream users can judge where the labels will generalize and where they will not. Write it alongside the data, because an annotation project already holds the guidelines, the label set, and the annotator demographics, and the sections it does not hold are easiest to write while the project is fresh.
A dataset is only as trustworthy as its documentation. You finish the annotation, export the labels, and hand off a file. Six months later someone trains a model on it, gets a strange result, and cannot tell whether the problem is their model or your data, because nobody wrote down who annotated it, how it was sampled, or what the labels were supposed to mean. The labels outlived the context that made them interpretable, and the documentation that would have prevented the confusion was cheap to write while the project was still fresh.
Irreproducibility and hidden bias in undocumented data
Undocumented data causes two problems, and both are expensive later. The first is irreproducibility. Without the sampling method, the guideline version, and the annotator pool, nobody can rebuild your dataset or explain a discrepancy against it, and the data becomes a black box that people either trust blindly or discard.
The second is hidden bias. A model trained on labels from a narrow annotator pool inherits that pool's blind spots, and if the pool was never documented, the bias is invisible until it surfaces in production. Documentation frameworks exist to prevent that outcome. They make the who and how of a dataset legible, so the biases can be seen before they ship.
Sections of a data statement
The data statement (Bender and Friedman, 2018) is the NLP-specific answer, a schema for characterizing a language dataset so users can understand how results might generalize and what biases the data carries. Six parts are worth writing down:
- Curation rationale. The rationale says what is in the dataset and why, including how items were sampled. A sample chosen from one subreddit is not the same dataset as a representative draw, and the rationale is where you say so.
- Language variety. Name the specific language and dialect, since "English" alone does not identify a variety. A model built on one variety may fail on another.
- Speaker demographics. Record who produced the source text.
- Annotator demographics. Record who produced the labels. On subjective tasks the annotator section is decisive, because the pool's composition shapes the labels, which is the whole argument for collecting demographics in the first place.
- Annotation guidelines. Include the instructions the labels were produced under, because the same label name means different things under different guidelines.
- Intended use. State what the dataset is for and, equally useful, what it is not for.
A data statement records who and how, so downstream users can judge where the labels generalize
Datasheets and model cards
Two neighboring frameworks complement data statements. Datasheets for datasets (Gebru et al.) borrow the idea from the electronics industry, where every component ships with a datasheet. Under the proposal, every dataset should ship with a document covering its motivation, composition, collection process, recommended uses, and maintenance. Datasheets are the general-purpose version and data statements the language-specific cousin, and the two overlap heavily.
Downstream of the data sits the model card (Mitchell et al., 2019), which documents a trained model's intended use and its performance broken down across demographic and other groups. The three form a chain. A datasheet or data statement documents the data, a model card documents what was built on it, and the annotator-demographics section of the first is what makes the group-wise evaluation in the last interpretable.
Release checklist for an annotation dataset
Before you release, confirm you can answer these six questions:
- How were items sampled, and from where?
- What language variety is this, and who wrote the source text?
- Who annotated it, how many people, and what is the demographic makeup of the pool?
- What guidelines did they follow, and which version?
- How was disagreement handled, aggregated to a gold label or kept as a distribution? Report agreement either way.
- What is this dataset for, and what should it not be used for?
A question you cannot answer marks a gap to close before the release goes out.
Data statement sections Potato already records
When you run annotation in Potato, the guideline, label, and annotator-demographics sections of the data statement already exist as project artifacts. The curation rationale, language variety, speaker demographics, and intended use are still yours to write.
The config is documentation. The YAML records the annotation schemes, the label sets, and the task structure, so the "what were the labels and how were they defined" part of a data statement is version-controlled alongside the data. The instructions and guidelines you wrote into the annotation guidelines are the guideline section, verbatim.
If you ran a prestudy phase, the demographics are already collected, and the annotator demographics are stored per annotator. Aggregated into distributions, never individual records, they become the annotator-demographics section, ready to paste in.
The export records who produced each label. Potato's export formats write the annotator's user_id on every record, so provenance travels with the data rather than getting stripped at the export step. A record from the JSONL exporter looks like this:
{
"instance_id": "doc_001",
"user_id": "user_1",
"labels": { "sentiment": { "positive": true } },
"spans": {},
"links": {}
}Timestamps are opt-in. Setting export_include_annotation_changes: true on a CSV or TSV export adds an annotation_changes file with one timestamped row per label change, so you can report when each label was set and revised.
When you publish to the Hugging Face Hub, Potato's huggingface exporter pushes a starter dataset card that lists each annotation scheme, its labels, and the number of annotation records. The card carries no curation rationale, guidelines, or demographics, so extend it from the config, the guidelines, and the prestudy demographics you already have, as the dataset card section of the exporting to HuggingFace walkthrough shows.
Further reading
These pages cover parts of a data statement in more depth.
- Collecting Annotator Demographics Responsibly, for the annotator-demographics section done right.
- Disagreement Is Signal, Not Noise, for documenting how you handled disagreement.
- Writing Effective Annotation Guidelines, which double as the guideline section of a data statement.
- Exporting Annotations for ML, for getting the labels and their metadata out cleanly.
References
Emily M. Bender & Batya Friedman (2018). Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science. Transactions of the Association for Computational Linguistics. https://aclanthology.org/Q18-1041/
Timnit Gebru et al. (2018). Datasheets for Datasets. arXiv preprint arXiv:1803.09010. https://arxiv.org/abs/1803.09010
Margaret Mitchell et al. (2019). Model Cards for Model Reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency. https://arxiv.org/abs/1810.03993