Who Cited It

SQuAD: 100,000+ Questions for Machine Comprehension of Text

2016 · 6,435 citations · 22 from inside this corpus

Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev low, Percy Liang

We present the Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles, where the answer to each question is a segment of text from the corresponding reading passage. We analyze the dataset to understand the types of reasoning required to answer the questions, leaning heavily on dependency and constituency trees. We build a strong logistic regression model, which achieves an F1 score of 51.0%, a significant improvement over a simple baseline (20%). However, human performance (86.8%) is much higher, indicating that the dataset presents a good challenge problem for future research. The dataset is freely available at https://stanford-qa.com

SQuAD: 100,000+ Questions for Machine Comprehension of Text (2016)SQuAD: 100,000+ Questions for…Building a Large Annotated Corpus of English: The Penn Treebank (1993)Building a Large Annotated Co…Building Watson: An Overview of the DeepQA Project (2010)Building Watson: An Overview …Teaching Machines to Read and Comprehend (2015)Teaching Machines to Read and…AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at S… (2018)AI-Assisted Pipeline for Dyna…BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019)BERT: Pre-training of Deep Bi…HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanit… (2019)HISTORIAE, History of Socio-C…On the Dangers of Stochastic Parrots (2021)On the Dangers of Stochastic …DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter (2019)DistilBERT, a distilled versi…ALBERT: A Lite BERT for Self-supervised Learning of Language\n Representations (2019)ALBERT: A Lite BERT for Self-…GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding (2018)GLUE: A Multi-Task Benchmark …Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Langu… (2022)Pre-train, Prompt, and Predic…Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (2019)Exploring the Limits of Trans…LXMERT: Learning Cross-Modality Encoder Representations from Transformers (2019)LXMERT: Learning Cross-Modali…On the Opportunities and Risks of Foundation Models (2021)On the Opportunities and Risk…Know What You Don’t Know: Unanswerable Questions for SQuAD (2018)Know What You Don’t Know: Una…Natural Questions: A Benchmark for Question Answering Research (2019)Natural Questions: A Benchmar…Language Models as Knowledge Bases? (2019)Language Models as Knowledge …Adversarial Examples: Attacks and Defenses for Deep Learning (2019)Adversarial Examples: Attacks…HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering (2018)HotpotQA: A Dataset for Diver…TinyBERT: Distilling BERT for Natural Language Understanding (2020)TinyBERT: Distilling BERT for…mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer (2021)mT5: A Massively Multilingual…Deep Learning--based Text Classification (2021)Deep Learning--based Text Cla…Pre-trained models for natural language processing: A survey (2020)Pre-trained models for natura…Reading Wikipedia to Answer Open-Domain Questions (2017)Reading Wikipedia to Answer O…ERNIE: Enhanced Language Representation with Informative Entities (2019)ERNIE: Enhanced Language Repr…
25 of 25 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

What cites it, inside the corpus

PaperYearCited
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at S…201846,036
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding201933,416
HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanit…201917,489
On the Dangers of Stochastic Parrots20216,426
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter20194,600
ALBERT: A Lite BERT for Self-supervised Learning of Language\n Representations20194,076
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding20184,055
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Langu…20223,836
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer20193,698
LXMERT: Learning Cross-Modality Encoder Representations from Transformers20192,335
On the Opportunities and Risks of Foundation Models20212,265
Know What You Don’t Know: Unanswerable Questions for SQuAD20182,196
Natural Questions: A Benchmark for Question Answering Research20192,095
Language Models as Knowledge Bases?20191,798
Adversarial Examples: Attacks and Defenses for Deep Learning20191,786
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering20181,743
TinyBERT: Distilling BERT for Natural Language Understanding20201,706
mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer20211,627
Deep Learning--based Text Classification20211,525
Pre-trained models for natural language processing: A survey20201,521
Reading Wikipedia to Answer Open-Domain Questions20171,454
ERNIE: Enhanced Language Representation with Informative Entities20191,435

Topics

Topic ModelingComputer Science
Natural Language Processing TechniquesComputer Science
Multimodal Machine Learning ApplicationsComputer Science

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports4 author record(s) attached.
  • supports27 reference(s) recorded.
  • neutralThe DOI carries no year to check against.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:43+00:00.

sha256 5cad55ac5d4d41e1…