Skip to content
Guides7 min read

Validated Survey Instruments for Annotation Studies: Personality, Affect, Wellbeing, and Demographics

When who annotates matters, use a validated questionnaire rather than one you invented. Potato ships 55 survey instruments for prestudy and poststudy phases.

Potato Team

When you want to measure something about your annotators, reach for a validated survey instrument such as the Big Five, PANAS, or a standard demographic battery before you write your own, because it comes with tested wordings, known reliability, and results comparable to a large body of prior work. Potato ships 55 of them, each usable in a prestudy or poststudy phase with a single config line.

Once you accept that the people doing your annotation shape the labels, the next question is what to measure about them. Age and education are the obvious start, but for subjective tasks the interesting predictors are often further afield, such as personality, values, mood on the day, and lived experience of the thing being judged. The temptation is to write a few quick questions and move on. Writing your own is usually a mistake, because a question you invent has no track record, no comparison group, and often a subtle wording flaw you will not notice until analysis.

Validated instruments versus homemade questions

A validated instrument is a questionnaire that researchers have tested for reliability, meaning it gives consistent results, and validity, meaning it measures what it claims to, usually across large samples and many studies. Borrowing one buys you three things a homemade question cannot: wordings that have been checked for ambiguity and bias, a scoring method with published norms, and comparability, because your numbers line up with everyone else who used the same instrument.

The cost of rolling your own shows up later. A gender question with the wrong options, a satisfaction scale that is subtly leading, or a personality question that half your annotators read differently each adds noise or bias you cannot separate from signal. The instrument authors already paid that cost so you do not have to.

Annotator characteristics that shape labels

Not every instrument belongs in every study, so match each one to a plausible effect on your task:

  • Demographics: who is annotating. The demographic batteries (ANES, GSS, ACS, and others) capture age, race, education, and the rest with standardized wordings. On offensiveness and politeness ratings, POPQUORN (Pei and Jurgens, 2023) found that annotators' backgrounds, including education, played a significant role in their judgments.
  • Personality and values: how someone judges. The BFI-2 (Soto and John, 2017), a 60-item Big Five inventory, and its ultra-brief cousin the Ten-Item Personality Inventory (Gosling et al., 2003) capture stable dispositions that can shape subjective ratings. The Moral Foundations Questionnaire (Graham et al., 2011) is a natural fit when the labels are moral judgments, since it measures the moral intuitions that drive them.
  • Affect: mood at labeling time. The PANAS (Watson et al., 1988) measures positive and negative affect. Run it in a poststudy phase and you can check whether mood tracked the ratings, which matters for emotionally loaded content.
  • Lived experience: standing to judge. The Everyday Discrimination Scale (Williams et al., 1997) measures day-to-day experience of discrimination. For tasks about offensiveness or hate directed at a group, whether an annotator has lived that is plausibly relevant to how they read it.
  • Wellbeing: protecting the annotator. Screeners like the PHQ-9 (Kroenke et al., 2001) and GAD-7 are not about the labels at all. On projects with harmful or distressing content, a light-touch wellbeing check helps you notice strain, provided you handle the responses with the care they demand.

Potato's survey-instrument library grouped into eight categories: demographic batteries, personality, mental health and wellbeing, affect and emotion, social and political attitudes, self-concept and social, response style, and short forms, with example instruments in each and the ones most relevant to annotation studies highlighted.The 55-instrument library, grouped by category, with the annotation-relevant ones highlighted

Sensitive screeners and survey length

Measuring your annotators is not free of risk, whatever you collect needs consent, and two of these categories carry real weight. Mental-health screeners are sensitive personal data. A PHQ-9 score is not a diagnosis, and it should never be treated as one or used to exclude someone from work. If you run one, say why, keep it optional, store it separately from anything identifying, and have a plan for what a concerning score means before you collect it. When in doubt, take it to your ethics board.

Length is its own tax. The Big Five Inventory-2 is 60 items, and a full battery stack can take longer than the annotation. Every extra question costs completion and attention, so lean on the short forms (e.g., the 10-item TIPI, the 2-item PHQ-2) unless you specifically need the long version, and cut anything you will not actually analyze. As with demographics, if there is no comparison you plan to run with a question, it does not go on the form.

Survey instruments in Potato

Potato includes a library of 55 validated instruments spanning personality, mental health, affect, social and political attitudes, and eight demographic batteries, all documented in Survey Instruments. You name these questionnaires rather than build them.

Reference one instrument by ID in a prestudy or poststudy phase:

yaml
phases:
  order: [consent, prestudy, annotation, poststudy]
 
  prestudy:
    type: prestudy
    instrument: "tipi"          # 10-item Big Five
 
  poststudy:
    type: poststudy
    instrument: "panas"         # affect, measured after the task

Stack several with instruments:, and append your own study-specific questions after a battery:

yaml
phases:
  prestudy:
    type: prestudy
    instruments:
      - "gss-demographics"      # standardized demographics
      - "srh"                   # single self-rated health item
    file: "surveys/study_specific.json"   # appended after the instruments

Each instrument carries its scoring metadata (method, reverse-coded items, range, and cutoffs), though Potato leaves the scoring to your analysis rather than computing it for you, which is the right call for anything clinical. The demographics-with-consent showcase puts the whole flow together, with a consent gate, a standardized demographic battery in the prestudy phase, and a subjective rating task, so the annotator background lands next to the labels where you can analyze it.

Further reading

References

Jiaxin Pei and David Jurgens (2023). When Do Annotator Demographics Matter? Measuring the Influence of Annotator Demographics with the POPQUORN Dataset. Proceedings of the 17th Linguistic Annotation Workshop (LAW-XVII). https://aclanthology.org/2023.law-1.25/

Christopher J. Soto and Oliver P. John (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology. https://doi.org/10.1037/pspp0000096

Samuel D. Gosling, Peter J. Rentfrow, and William B. Swann (2003). A very brief measure of the Big-Five personality domains. Journal of Research in Personality. https://doi.org/10.1016/S0092-6566(03)00046-1

Jesse Graham et al. (2011). Mapping the moral domain. Journal of Personality and Social Psychology. https://doi.org/10.1037/a0021847

David Watson, Lee Anna Clark, and Auke Tellegen (1988). Development and validation of brief measures of positive and negative affect: The PANAS scales. Journal of Personality and Social Psychology. https://doi.org/10.1037/0022-3514.54.6.1063

David R. Williams et al. (1997). Racial Differences in Physical and Mental Health: Socio-economic Status, Stress and Discrimination. Journal of Health Psychology. https://doi.org/10.1177/135910539700200305

Kurt Kroenke, Robert L. Spitzer, and Janet B. W. Williams (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine. https://doi.org/10.1046/j.1525-1497.2001.016009606.x