Skip to content
Guides7 min read

Impostare l'annotazione per la moderazione dei contenuti

Come configurare Potato per il rilevamento della tossicità, la classificazione dell'incitamento all'odio e l'etichettatura di contenuti sensibili, tenendo conto del benessere degli annotatori.

Potato Team

Etichettare contenuti tossici non è come etichettare recensioni di film. Il lavoro logora chi lo fa, le linee guida restano sempre un po' sfocate e i casi che contano di più sono di solito quelli su cui gli annotatori non sono d'accordo. Questa guida spiega come impostare in Potato un task di moderazione dei contenuti che prenda sul serio questi problemi, partendo dalle persone che il lavoro lo fanno.

Benessere degli annotatori

Leggere per ore contenuti d'odio e immagini crude ha effetti concreti su chi lo fa. Alcune impostazioni aiutano a rendere il lavoro sostenibile.

Configurazione del benessere

yaml
wellbeing:
  # Content warnings
  warnings:
    enabled: true
    show_before_session: true
    message: |
      This task involves reviewing potentially offensive content including
      hate speech, harassment, and explicit material. Take breaks as needed.
 
  # Break reminders
  breaks:
    enabled: true
    reminder_interval: 30  # minutes
    break_duration: 5  # suggested minutes
    message: "Consider taking a short break. Your wellbeing matters."
 
  # Session limits
  limits:
    max_session_duration: 120  # minutes
    max_items_per_session: 100
    cooldown_between_sessions: 60  # minutes
 
  # Easy exit
  exit:
    allow_immediate_exit: true
    no_penalty_exit: true
    exit_button_prominent: true
    exit_message: "No problem. Take care of yourself."
 
  # Support resources
  resources:
    show_support_link: true
    support_url: "https://yourorg.com/support"
    hotline_number: "1-800-XXX-XXXX"

Sfocatura dei contenuti

yaml
display:
  # Blur images by default
  image_display:
    blur_by_default: true
    blur_amount: 20
    click_to_reveal: true
    reveal_duration: 10  # auto-blur after 10 seconds
 
  # Text content warnings
  text_display:
    show_severity_indicator: true
    expandable_content: true
    default_collapsed: true

Classificazione della tossicità

Tossicità su più livelli

yaml
annotation_schemes:
  - annotation_type: radio
    name: toxicity_level
    description: "Rate the toxicity level of this content"
    labels:
      - name: none
        label: "Not Toxic"
        description: "No harmful content"
 
      - name: mild
        label: "Mildly Toxic"
        description: "Rude or insensitive but not severe"
 
      - name: moderate
        label: "Moderately Toxic"
        description: "Clearly offensive or harmful"
 
      - name: severe
        label: "Severely Toxic"
        description: "Extremely offensive, threatening, or dangerous"

Categorie di tossicità

yaml
annotation_schemes:
  - annotation_type: multiselect
    name: toxicity_types
    description: "Select all types of toxicity present"
    labels:
      - name: profanity
        label: "Profanity/Obscenity"
        description: "Swear words, vulgar language"
 
      - name: insult
        label: "Insults"
        description: "Personal attacks, name-calling"
 
      - name: threat
        label: "Threats"
        description: "Threats of violence or harm"
 
      - name: hate_speech
        label: "Hate Speech"
        description: "Targeting protected groups"
 
      - name: harassment
        label: "Harassment"
        description: "Targeted, persistent hostility"
 
      - name: sexual
        label: "Sexual Content"
        description: "Explicit or suggestive content"
 
      - name: self_harm
        label: "Self-Harm/Suicide"
        description: "Promoting or glorifying self-harm"
 
      - name: misinformation
        label: "Misinformation"
        description: "Demonstrably false claims"
 
      - name: spam
        label: "Spam/Scam"
        description: "Unwanted promotional content"

Rilevamento dell'incitamento all'odio

Gruppi presi di mira

yaml
annotation_schemes:
  - annotation_type: multiselect
    name: target_groups
    description: "Which groups are targeted? (if hate speech detected)"
 
    labels:
      - name: race_ethnicity
        label: "Race/Ethnicity"
 
      - name: religion
        label: "Religion"
 
      - name: gender
        label: "Gender"
 
      - name: sexual_orientation
        label: "Sexual Orientation"
 
      - name: disability
        label: "Disability"
 
      - name: nationality
        label: "Nationality/Origin"
 
      - name: age
        label: "Age"
 
      - name: other
        label: "Other Protected Group"

Gravità dell'incitamento all'odio

yaml
annotation_schemes:
  - annotation_type: radio
    name: hate_severity
    description: "Severity of hate speech"
 
    labels:
      - name: implicit
        label: "Implicit"
        description: "Coded language, dog whistles"
 
      - name: explicit_mild
        label: "Explicit - Mild"
        description: "Clear but not threatening"
 
      - name: explicit_severe
        label: "Explicit - Severe"
        description: "Dehumanizing, threatening, or violent"

Moderazione in base al contesto

Regole specifiche per piattaforma

yaml
# Context affects what's acceptable
annotation_schemes:
  - annotation_type: radio
    name: context_appropriate
    description: "Is this content appropriate for the platform context?"
 
    labels:
      - name: appropriate
        label: "Appropriate for Context"
 
      - name: borderline
        label: "Borderline"
 
      - name: inappropriate
        label: "Inappropriate for Context"
 
  - annotation_type: text
    name: context_notes
    description: "Explain your contextual reasoning"

Rilevamento dell'intenzione

yaml
annotation_schemes:
  - annotation_type: radio
    name: intent
    description: "What is the apparent intent?"
    labels:
      - name: genuine_attack
        label: "Genuine Attack"
        description: "Intent to harm or offend"
 
      - name: satire
        label: "Satire/Parody"
        description: "Mocking toxic behavior"
 
      - name: quote
        label: "Quote/Report"
        description: "Reporting or discussing toxic content"
 
      - name: reclaimed
        label: "Reclaimed Language"
        description: "In-group use of slurs"
 
      - name: unclear
        label: "Unclear Intent"

Moderazione delle immagini

Classificazione dei contenuti visivi

yaml
annotation_schemes:
  - annotation_type: multiselect
    name: image_violations
    description: "Select all policy violations"
    labels:
      - name: nudity
        label: "Nudity/Sexual Content"
 
      - name: violence_graphic
        label: "Graphic Violence"
 
      - name: gore
        label: "Gore/Disturbing Content"
 
      - name: hate_symbols
        label: "Hate Symbols"
 
      - name: dangerous_acts
        label: "Dangerous Acts"
 
      - name: child_safety
        label: "Child Safety Concern"
        priority: critical
        escalate: true
 
      - name: none
        label: "No Violations"
 
  - annotation_type: radio
    name: action_recommendation
    description: "Recommended action"
    labels:
      - name: approve
        label: "Approve"
      - name: age_restrict
        label: "Age-Restrict"
      - name: warning_label
        label: "Add Warning Label"
      - name: remove
        label: "Remove"
      - name: escalate
        label: "Escalate to Specialist"

Controllo di qualità

yaml
quality_control:
  # Calibration for subjective content
  calibration:
    enabled: true
    frequency: 20  # Every 20 items
    items: calibration/moderation_gold.json
    feedback: true
    recalibrate_on_drift: true
 
  # High redundancy for borderline cases
  redundancy:
    annotations_per_item: 3
    increase_for_borderline: 5
    agreement_threshold: 0.67
 
  # Expert escalation
  escalation:
    enabled: true
    triggers:
      - field: toxicity_level
        value: severe
      - field: image_violations
        contains: child_safety
    escalate_to: trust_safety_team
 
  # Distribution monitoring
  monitoring:
    track_distribution: true
    alert_on_skew: true
    expected_distribution:
      none: 0.4
      mild: 0.3
      moderate: 0.2
      severe: 0.1

Configurazione completa

yaml
annotation_task_name: "Content Moderation"
 
# Wellbeing first
wellbeing:
  warnings:
    enabled: true
    message: "This task contains potentially offensive content."
  breaks:
    reminder_interval: 30
    message: "Remember to take breaks."
  limits:
    max_session_duration: 90
    max_items_per_session: 75
 
display:
  # Blur sensitive content
  image_display:
    blur_by_default: true
    click_to_reveal: true
 
  # Show platform context
  metadata_display:
    show_fields: [platform, community, report_reason]
 
annotation_schemes:
  # Toxicity level
  - annotation_type: radio
    name: toxicity
    description: "Toxicity level"
    labels:
      - name: none
        label: "None"
      - name: mild
        label: "Mild"
      - name: moderate
        label: "Moderate"
      - name: severe
        label: "Severe"
 
  # Categories
  - annotation_type: multiselect
    name: categories
    description: "Types of harmful content (select all)"
    labels:
      - name: hate
        label: "Hate Speech"
      - name: harassment
        label: "Harassment"
      - name: violence
        label: "Violence/Threats"
      - name: sexual
        label: "Sexual Content"
      - name: self_harm
        label: "Self-Harm"
      - name: spam
        label: "Spam"
      - name: none
        label: "None"
 
  # Confidence
  - annotation_type: likert
    name: confidence
    description: "How confident are you?"
    size: 5
    min_label: "Uncertain"
    max_label: "Very Confident"
 
  # Notes
  - annotation_type: text
    name: notes
    description: "Additional notes (optional)"
    rows: 4
 
quality_control:
  redundancy:
  calibration:
  escalation:
 
output_annotation_dir: annotations/
export_annotation_format: jsonl

Scrivere le linee guida

Gran parte del disaccordo tra annotatori nasce da linee guida vaghe, quindi è qui che lo sforzo rende. Spiega che cosa distingue «lieve» da «moderato» invece di lasciare che gli annotatori tirino a indovinare. Mostra i casi limite, con la spiegazione del perché ciascuno è finito dov'è finito. Di' come piattaforma e pubblico cambiano il giudizio, dato che la stessa frase può andare bene in una community ed essere una violazione in un'altra. E spiega agli annotatori che cosa fare quando l'intenzione è davvero poco chiara, invece di far finta che non succeda mai. Metti in conto di dover rivedere tutto questo man mano che compaiono contenuti di tipo nuovo.

Sostenere gli annotatori

Non lasciare la stessa persona sui contenuti tossici tutto il giorno: fai ruotare il lavoro. Tieni le risorse di supporto psicologico a un clic di distanza, dai alle persone un canale reale per segnalare problemi e riconosci che è un lavoro pesante. Lo è.

Per il funzionamento degli schemi di classificazione sottostanti, vedi la documentazione sull'annotazione di testo. La guida al controllo di qualità tratta ridondanza e calibrazione più in dettaglio.


Documentazione completa su Schemi di Annotazione.