PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 20182,198 citationsOpen Access

Know What You Don’t Know: Unanswerable Questions for SQuAD

PRPranav RajpurkarRJRobin JiaPLPercy Liang

Key Points

  • This research aims to enhance extractive reading comprehension systems by introducing a dataset that includes both answerable and adversarial unanswerable questions.
  • Introduced SQUADRUN, a dataset with over 50,000 adversarial unanswerable questions created by crowdworkers.
  • Combined with the existing Stanford Question Answering Dataset (SQuAD).
  • Evaluated performance of models based on F1 scores on both SQuAD and SQUADRUN.
  • A strong neural system scored 86% F1 on SQuAD but only 66% F1 on SQUADRUN, highlighting performance disparity.
  • Performance drop indicates challenges for models in distinguishing between answerable and unanswerable questions.

Abstract

Extractive reading comprehension systems can often locate the correct answer to a question in a context document, but they also tend to make unreliable guesses on questions for which the correct answer is not stated in the context. Existing datasets either focus exclusively on answerable questions, or use automatically generated unanswerable questions that are easy to identify. To address these weaknesses, we present SQUADRUN, a new dataset that combines the existing Stanford Question Answering Dataset (SQuAD) with over 50,000 unanswerable questions written adversarially by crowdworkers to look similar to answerable ones. To do well on SQUADRUN, systems must not only answer questions when possible, but also determine when no answer is supported by the paragraph and abstain from answering. SQUADRUN is a challenging natural language understanding task for existing models: a strong neural system that gets 86% F1 on SQuAD achieves only 66% F1 on SQUADRUN. We release SQUADRUN to the community as the successor to SQuAD.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rajpurkar et al. (2018) studied this question.

synapsesocial.com/papers/69dd6697c5e71f7918100f24https://doi.org/10.18653/v1/p18-2124
Ask AI
Helpful
Bookmark
Share
View Full Paper