Who Cited It

Text Data Augmentation for Deep Learning

2021 · Journal Of Big Data · 1,701 citations · 0 from inside this corpus

Connor Shorten, Taghi M. Khoshgoftaar, Borko Furht

Natural Language Processing (NLP) is one of the most captivating applications of Deep Learning. In this survey, we consider how the Data Augmentation training strategy can aid in its development. We begin with the major motifs of Data Augmentation summarized into strengthening local decision boundaries, brute force training, causality and counterfactual examples, and the distinction between meaning and form. We follow these motifs with a concrete list of augmentation frameworks that have been developed for text data. Deep Learning generally struggles with the measurement of generalization and characterization of overfitting. We highlight studies that cover how augmentations can construct test sets for generalization. NLP is at an early stage in applying Data Augmentation compared to Computer Vision. We highlight the key differences and promising ideas that have yet to be tested in NLP. For the sake of practical implementation, we describe tools that facilitate Data Augmentation such as the use of consistency regularization, controllers, and offline and online augmentation pipelines, to preview a few. Finally, we discuss interesting topics around Data Augmentation in NLP such as task-specific augmentations, the use of prior knowledge in self-supervised learning versus Data Augmentation, intersections with transfer and multi-task learning, and ideas for AI-GAs (AI-Generating Algorithms). We hope this paper inspires further research interest in Text Data Augmentation.

Text Data Augmentation for Deep Learning (2021)Text Data Augmentation for De…Dropout: a simple way to prevent neural networks from overfitting (2014)Dropout: a simple way to prev…BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019)BERT: Pre-training of Deep Bi…SMOTE: Synthetic Minority Over-sampling Technique (2002)SMOTE: Synthetic Minority Ove…WordNet (1995)WordNetDistilling the Knowledge in a Neural Network (2015)Distilling the Knowledge in a…Momentum Contrast for Unsupervised Visual Representation Learning (2020)Momentum Contrast for Unsuper…Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (2019)Sentence-BERT: Sentence Embed…Building a Large Annotated Corpus of English: The Penn Treebank (1993)Building a Large Annotated Co…A Simple Framework for Contrastive Learning of Visual Representations (2020)A Simple Framework for Contra…A survey of transfer learning (2016)A survey of transfer learningAdvances and Open Problems in Federated Learning (2020)Advances and Open Problems in…Causality: Models, Reasoning and Inference (2001)Causality: Models, Reasoning …DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter (2019)DistilBERT, a distilled versi…GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding (2018)GLUE: A Multi-Task Benchmark …Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Met… (2026)Affordance-Compiled Intellige…Survey on deep learning with class imbalance (2019)Survey on deep learning with …Causality: models, reasoning, and inference (2000)Causality: models, reasoning,…A Survey on Knowledge Graphs: Representation, Acquisition, and Applications (2021)A Survey on Knowledge Graphs:…EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Ta… (2019)RUSBoost: A Hybrid Approach to Alleviating Class Imbalance (2009)RUSBoost: A Hybrid Approach t…GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding (2018)GLUE: A Multi-Task Benchmark …AutoML: A survey of the state-of-the-art (2020)AutoML: A survey of the state…Unsupervised Data Augmentation for Consistency Training (2019)Unsupervised Data Augmentatio…
23 of 23 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

PaperYearCited
Dropout: a simple way to prevent neural networks from overfitting201434,236
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding201933,416
SMOTE: Synthetic Minority Over-sampling Technique200232,402
WordNet199514,235
Distilling the Knowledge in a Neural Network201514,099
Momentum Contrast for Unsupervised Visual Representation Learning202012,515
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks201911,789
Building a Large Annotated Corpus of English: The Penn Treebank19937,538
A Simple Framework for Contrastive Learning of Visual Representations20207,341
A survey of transfer learning20166,290
Advances and Open Problems in Federated Learning20205,383
Causality: Models, Reasoning and Inference20014,831
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter20194,600
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding20184,055
Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Met…20263,059
Survey on deep learning with class imbalance20192,971
Causality: models, reasoning, and inference20002,920
A Survey on Knowledge Graphs: Representation, Acquisition, and Applications20212,862
EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Ta…20191,952
RUSBoost: A Hybrid Approach to Alleviating Class Imbalance20091,890
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding20181,861
AutoML: A survey of the state-of-the-art20201,705
Unsupervised Data Augmentation for Consistency Training20191,624

Topics

Topic ModelingComputer Science
Text and Document Classification TechnologiesComputer Science
Machine Learning in HealthcareComputer Science

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports3 author record(s) attached.
  • supports129 reference(s) recorded.
  • neutralThe DOI carries no year to check against.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:56+00:00.

sha256 7bc26169e93749d1…