Natural Questions: A Benchmark for Question Answering Research
Tom Kwiatkowski low, Jennimaria Palomaki low, Olivia Redfield low, Michael Collins, Ankur P. Parikh, Chris Alberti low, Danielle Epstein low, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey low, Ming‐Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc V. Le, Slav Petrov
We present the Natural Questions corpus, a question answering data set. Questions consist of real anonymized, aggregated queries issued to the Google search engine. An annotator is presented with a question along with a Wikipedia page from the top 5 search results, and annotates a long answer (typically a paragraph) and a short answer (one or more entities) if present on the page, or marks null if no long/short answer is present. The public release consists of 307,373 training examples with single annotations; 7,830 examples with 5-way annotations for development data; and a further 7,842 examples with 5-way annotated sequestered as test data. We present experiments validating quality of the data. We also describe analysis of 25-way annotations on 302 examples, giving insights into human variability on the annotation task. We introduce robust metrics for the purposes of evaluating question answering systems; demonstrate high human upper bounds on these metrics; and establish baseline results using competitive methods drawn from related literature.
What this paper cites, inside the corpus
What cites it, inside the corpus
Links
Topics
| Topic Modeling | Computer Science |
| Multimodal Machine Learning Applications | Computer Science |
| Natural Language Processing Techniques | Computer Science |
Is this record sound?
complete
Nothing in this record contradicts itself and no field we check is missing.
- supports18 author record(s) attached.
- supports34 reference(s) recorded.
- neutralThe DOI carries no year to check against.
- supportsA title is present.
Provenance
sha256 6e314d9c6052d1ed…