Who Cited It

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

2019 · 33,416 citations · 39 from inside this corpus

Jacob Devlin, Ming‐Wei Chang, Kenton Lee, Kristina Toutanova

Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019)BERT: Pre-training of Deep Bi…Glove: Global Vectors for Word Representation (2014)Glove: Global Vectors for Wor…Distributed Representations of Words and Phrases and their Compositionality (2013)Distributed Representations o…Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank (2013)Recursive Deep Models for Sem…SQuAD: 100,000+ Questions for Machine Comprehension of Text (2016)SQuAD: 100,000+ Questions for…A unified architecture for natural language processing (2008)A unified architecture for na…Distributed Representations of Sentences and Documents (2014)Distributed Representations o…Class-based n -gram models of natural language (1992)Class-based n -gram models of…Supervised Learning of Universal Sentence Representations from Natural\n Language Inferen… (2017)Supervised Learning of Univer…Word Representations: A Simple and General Method for Semi-Supervised Learning (2010)Word Representations: A Simpl…Domain adaptation with structural correspondence learning (2006)Domain adaptation with struct…A Decomposable Attention Model for Natural Language Inference (2016)A Decomposable Attention Mode…A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data (2005)A Framework for Learning Pred…HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanit… (2019)HISTORIAE, History of Socio-C…Momentum Contrast for Unsupervised Visual Representation Learning (2020)Momentum Contrast for Unsuper…Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (2019)Sentence-BERT: Sentence Embed…Graph neural networks: A review of methods and applications (2020)Graph neural networks: A revi…Emerging Properties in Self-Supervised Vision Transformers (2021)DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter (2019)DistilBERT, a distilled versi…Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Langu… (2022)Pre-train, Prompt, and Predic…Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (2019)Exploring the Limits of Trans…Transformer-XL: Attentive Language Models beyond a Fixed-Length Context (2019)Transformer-XL: Attentive Lan…SciBERT: A Pretrained Language Model for Scientific Text (2019)SciBERT: A Pretrained Languag…Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Met… (2026)Affordance-Compiled Intellige…Language Models are Few-Shot Learners (2020)Language Models are Few-Shot …SimCSE: Simple Contrastive Learning of Sentence Embeddings (2021)SimCSE: Simple Contrastive Le…CodeBERT: A Pre-Trained Model for Programming and Natural Languages (2020)CodeBERT: A Pre-Trained Model…LoRA Fine-Tuning of a 3B Code LLM for Algorithmic Efficiency (2021)LoRA Fine-Tuning of a 3B Code…LXMERT: Learning Cross-Modality Encoder Representations from Transformers (2019)LXMERT: Learning Cross-Modali…Prefix-Tuning: Optimizing Continuous Prompts for Generation (2021)Prefix-Tuning: Optimizing Con…On the Opportunities and Risks of Foundation Models (2021)On the Opportunities and Risk…Longformer: The Long-Document Transformer (2020)Longformer: The Long-Document…BERTScore: Evaluating Text Generation with BERT (2019)BERTScore: Evaluating Text Ge…Shortcut learning in deep neural networks (2020)A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and… (2024)A Survey on Hallucination in …EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Ta… (2019)EDA: Easy Data Augmentation T…XLNet: Generalized Autoregressive Pretraining for Language Understanding (2019)
36 of 48 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

What cites it, inside the corpus

PaperYearCited
HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanit…201917,489
Momentum Contrast for Unsupervised Visual Representation Learning202012,515
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks201911,789
Graph neural networks: A review of methods and applications20205,808
Emerging Properties in Self-Supervised Vision Transformers20215,394
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter20194,600
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Langu…20223,836
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer20193,698
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context20193,202
SciBERT: A Pretrained Language Model for Scientific Text20193,109
Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Met…20263,059
Language Models are Few-Shot Learners20203,020
SimCSE: Simple Contrastive Learning of Sentence Embeddings20212,645
CodeBERT: A Pre-Trained Model for Programming and Natural Languages20202,630
LoRA Fine-Tuning of a 3B Code LLM for Algorithmic Efficiency20212,531
LXMERT: Learning Cross-Modality Encoder Representations from Transformers20192,335
Prefix-Tuning: Optimizing Continuous Prompts for Generation20212,322
On the Opportunities and Risks of Foundation Models20212,265
Longformer: The Long-Document Transformer20202,206
BERTScore: Evaluating Text Generation with BERT20192,065
Shortcut learning in deep neural networks20202,040
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and…20241,986
EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Ta…20191,952
XLNet: Generalized Autoregressive Pretraining for Language Understanding20191,854

Topics

Topic ModelingComputer Science
Natural Language Processing TechniquesComputer Science
Computational and Text Analysis MethodsSocial Sciences

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports4 author record(s) attached.
  • supports52 reference(s) recorded.
  • neutralThe DOI carries no year to check against.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:40+00:00.

sha256 7e3d99a592f7f61f…