Who Cited It

TinyBERT: Distilling BERT for Natural Language Understanding

2020 · 1,706 citations · 3 from inside this corpus

Xiaoqi Jiao low, Yichun Yin low, Lifeng Shang low, Xin Jiang, Xiao Dong Chen, Linlin Li, Fang Wang

Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually computationally expensive, so it is difficult to efficiently execute them on resourcerestricted devices. To accelerate inference and reduce model size while maintaining accuracy, we first propose a novel Transformer distillation method that is specially designed for knowledge distillation (KD) of the Transformer-based models. By leveraging this new KD method, the plenty of knowledge encoded in a large "teacher" BERT can be effectively transferred to a small "student" Tiny-BERT. Then, we introduce a new two-stage learning framework for TinyBERT, which performs Transformer distillation at both the pretraining and task-specific learning stages. This framework ensures that TinyBERT can capture the general-domain as well as the task-specific knowledge in BERT.

TinyBERT: Distilling BERT for Natural Language Understanding (2020)TinyBERT: Distilling BERT for…Exploiting Generative AI to Scale up Intelligent Tutoring Systems (2023)Exploiting Generative AI to S…AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at S… (2018)AI-Assisted Pipeline for Dyna…Glove: Global Vectors for Word Representation (2014)Glove: Global Vectors for Wor…HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanit… (2019)HISTORIAE, History of Socio-C…Distilling the Knowledge in a Neural Network (2015)Distilling the Knowledge in a…Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer (2019)Exploring the Limits of Trans…Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank (2013)Recursive Deep Models for Sem…SQuAD: 100,000+ Questions for Machine Comprehension of Text (2016)SQuAD: 100,000+ Questions for…DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter (2019)DistilBERT, a distilled versi…ALBERT: A Lite BERT for Self-supervised Learning of Language\n Representations (2019)ALBERT: A Lite BERT for Self-…GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding (2018)GLUE: A Multi-Task Benchmark …Know What You Don’t Know: Unanswerable Questions for SQuAD (2018)Know What You Don’t Know: Una…Question Answering For Toxicological Information Extraction (2022)Question Answering For Toxico…Language Models are Few-Shot Learners (2020)Language Models are Few-Shot …Pre-trained models for natural language processing: A survey (2020)Pre-trained models for natura…A Brief Overview of ChatGPT: The History, Status Quo and Potential Future Development (2023)A Brief Overview of ChatGPT: …
16 of 16 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

What cites it, inside the corpus

Topics

Topic ModelingComputer Science
Natural Language Processing TechniquesComputer Science
Multimodal Machine Learning ApplicationsComputer Science

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports7 author record(s) attached.
  • supports46 reference(s) recorded.
  • supportsThe DOI's year agrees with the publication year.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:53+00:00.

sha256 db1645b78a57e29a…