Skip to content
intermediatetext

WNUT-2017: Emerging and Rare Entity Recognition

WNUT-2017 is a named entity recognition benchmark for novel and rare entities in noisy social media text, organized at W-NUT@EMNLP 2017. This Potato config reproduces its 6-type span annotation task.

About this dataset

WNUT-2017 was the shared task on Novel and Emerging Entity Recognition, organized by Leon Derczynski, Eric Nichols, Marieke van Erp, and Nut Limsopatham at the 3rd Workshop on Noisy User-generated Text (W-NUT@EMNLP 2017) in Copenhagen. It targets entities that are rare or unseen at training time, the hardest case for NER systems trained on standard newswire.

The corpus draws each split from a different platform to force generalization: training from Twitter (3,395 posts, 62,729 tokens), development from YouTube comments (1,009 posts, 15,733 tokens), and test from Reddit and StackExchange (1,287 posts, 23,394 tokens). Named-entity tokens make up roughly 5 to 8 percent of each split.

Annotators mark spans in one of six types: person, location, corporation, product, creative-work, and group. Systems are scored on two F1 measures: a standard entity-level F1 over all mentions, plus a surface-form F1 computed over unique entity strings, which rewards recognizing a diverse range of entities rather than only frequent ones. The best system reached 41.86 entity F1 and 40.24 surface-form F1.

The Potato config below reproduces the task with a span scheme over the six entity types and a follow-up radio question on entity composition, so annotators highlight each mention and label the text in noisy social media data.

Entity types
6 (person, location, corporation, product, creative-work, group)
Training set
3,395 Twitter posts, 62,729 tokens
Dev set
1,009 YouTube comment posts, 15,733 tokens
Test set
1,287 Reddit + StackExchange posts, 23,394 tokens
Metrics
Entity F1 and surface-form F1
Best system
41.86 entity F1, 40.24 surface-form F1
PERORGLOCPERORGLOCDATESelect text to annotate

Configuration Fileconfig.yaml

This Potato config reproduces the annotation task. Save it as config.yaml and run potato start config.yaml to try it.

yaml
# WNUT-2017 - Emerging and Rare Entity Recognition
# Based on Derczynski et al., W-NUT@EMNLP 2017
# Paper: https://aclanthology.org/W17-4418/
# Dataset: https://noisy-text.github.io/2017/emerging-rare-entities.html
#
# This task presents social media text (tweets, Reddit posts) for
# named entity recognition with a focus on emerging and rare entities.
# Annotators highlight entity spans and classify the overall entity
# composition of the text.
#
# Entity Types:
# - Person: Names of people
# - Location: Places, geographic locations
# - Corporation: Companies and organizations
# - Product: Commercial products, software, services
# - Creative Work: Movies, books, songs, games, etc.
# - Group: Sports teams, bands, political organizations
#
# Annotation Guidelines:
# 1. Read the social media text carefully
# 2. Highlight all entity mentions using the appropriate entity type
# 3. Include the full entity name (e.g., "New York City" not just "New York")
# 4. Classify whether the text contains novel, standard, or no entities

annotation_task_name: "WNUT-2017 - Emerging and Rare Entity Recognition"
task_dir: "."

data_files:
  - sample-data.json

item_properties:
  id_key: "id"
  text_key: "text"

output_annotation_dir: "annotation_output/"
output_annotation_format: "json"

port: 8000
server_name: localhost

annotation_schemes:
  # Step 1: Highlight entity spans
  - annotation_type: span
    name: entity_spans
    description: "Highlight all named entities in the text"
    labels:
      - "Person"
      - "Location"
      - "Corporation"
      - "Product"
      - "Creative Work"
      - "Group"
    label_colors:
      "Person": "#3b82f6"
      "Location": "#ef4444"
      "Corporation": "#22c55e"
      "Product": "#f59e0b"
      "Creative Work": "#8b5cf6"
      "Group": "#06b6d4"
    keyboard_shortcuts:
      "Person": "1"
      "Location": "2"
      "Corporation": "3"
      "Product": "4"
      "Creative Work": "5"
      "Group": "6"

  # Step 2: Entity composition classification
  - annotation_type: radio
    name: entity_composition
    description: "What type of entities does this text contain?"
    labels:
      - "Contains Novel Entities"
      - "Standard Entities Only"
      - "No Entities"
    keyboard_shortcuts:
      "Contains Novel Entities": "7"
      "Standard Entities Only": "8"
      "No Entities": "9"
    tooltips:
      "Contains Novel Entities": "Text mentions emerging, recently created, or rarely seen entities"
      "Standard Entities Only": "Text mentions only well-known, established entities"
      "No Entities": "Text contains no named entities"

annotation_instructions: |
  You will be shown a social media post (from Twitter or Reddit). Your task is to:
  1. Highlight all named entity mentions in the text using the appropriate category.
  2. Classify whether the text contains novel/emerging entities, standard entities, or no entities.

  Entity categories:
  - Person: Names of individuals
  - Location: Geographic places, cities, countries
  - Corporation: Companies, organizations, institutions
  - Product: Software, devices, commercial products
  - Creative Work: Movies, books, songs, TV shows, games
  - Group: Sports teams, bands, political groups

  Pay special attention to emerging or novel entities that may not be well-known.

html_layout: |
  <div style="padding: 15px; max-width: 800px; margin: auto;">
    <div style="background: #f0f9ff; border: 1px solid #bae6fd; border-radius: 8px; padding: 16px; margin-bottom: 16px;">
      <strong style="color: #0369a1;">Social Media Text:</strong>
      <p style="font-size: 16px; line-height: 1.7; margin: 8px 0 0 0;">{{text}}</p>
    </div>
  </div>

allow_all_users: true
instances_per_annotator: 50
annotation_per_instance: 2
allow_skip: true
skip_reason_required: false

Sample Datasample-data.json

json
[
  {
    "id": "wnut_001",
    "text": "Just saw the new Deadpool movie with @john_smith at AMC Theatres downtown. Absolutely hilarious!"
  },
  {
    "id": "wnut_002",
    "text": "Anyone tried the new Pixel 9 Pro? Thinking of switching from my Samsung Galaxy. Google really stepped up their game."
  }
]

// ... and 8 more items

Get This Design

View on GitHub

Clone or download from the repository

Quick start:

git clone https://github.com/davidjurgens/potato-showcase.git
cd potato-showcase/text/named-entity-recognition/wnut2017-emerging-entities
potato start config.yaml

Dataset & paper

Derczynski et al., W-NUT@EMNLP 2017

Citation (BibTeX)

bibtex
@inproceedings{derczynski-etal-2017-results,
    title = "Results of the {WNUT}2017 Shared Task on Novel and Emerging Entity Recognition",
    author = "Derczynski, Leon  and Nichols, Eric  and van Erp, Marieke  and Limsopatham, Nut",
    booktitle = "Proceedings of the 3rd Workshop on Noisy User-generated Text",
    month = sep,
    year = "2017",
    address = "Copenhagen, Denmark",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/W17-4418",
    pages = "140--147"
}

Details

Annotation Types

spanradio

Domain

NLPSocial Media

Use Cases

Named Entity RecognitionInformation ExtractionSocial Media Analysis

Tags

neremerging-entitiessocial-mediatwitterredditwnut2017

Found an issue or want to improve this design?

Open an Issue