Who Cited It

Transformer-XL: Attentive Language Models beyond a Fixed-Length Context

2019 · 3,202 citations · 9 from inside this corpus

Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, Ruslan Salakhutdinov low

Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence. It consists of a segment-level recurrence mechanism and a novel positional encoding scheme. Our method not only enables capturing longer-term dependency, but also resolves the context fragmentation problem. As a result, Transformer-XL learns dependency that is 80% longer than RNNs and 450% longer than vanilla Transformers, achieves better performance on both short and long sequences, and is up to 1,800+ times faster than vanilla Transformers during evaluation. Notably, we improve the state-ofthe-art results of bpc/perplexity to 0.99 on en-wiki8, 1.08 on text8, 18.3 on WikiText-103, 21.8 on One Billion Word, and 54.5 on Penn Treebank (without finetuning). When trained only on WikiText-103, Transformer-XL manages to generate reasonably coherent, novel text articles with thousands of tokens. Our code, pretrained models, and hyperparameters are available in both Tensorflow and PyTorch 1 .

Transformer-XL: Attentive Language Models beyond a Fixed-Length Context (2019)Transformer-XL: Attentive Lan…Long Short-Term Memory (1997)Long Short-Term MemoryExploiting Generative AI to Scale up Intelligent Tutoring Systems (2023)Exploiting Generative AI to S…AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at S… (2018)AI-Assisted Pipeline for Dyna…BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019)BERT: Pre-training of Deep Bi…Neural Machine Translation by Jointly Learning to Align and Translate (2014)Neural Machine Translation by…Recurrent neural network based language model (2010)Recurrent neural network base…An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Mode… (2018)An Empirical Evaluation of Ge…Neural Architecture Search with Reinforcement Learning (2016)Neural Architecture Search wi…Transformers: State-of-the-Art Natural Language Processing (2020)Transformers: State-of-the-Ar…Conformer: Convolution-augmented Transformer for Speech Recognition (2020)Conformer: Convolution-augmen…ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning (2021)ProtTrans: Toward Understandi…On the Opportunities and Risks of Foundation Models (2021)On the Opportunities and Risk…Longformer: The Long-Document Transformer (2020)Longformer: The Long-Document…Natural language processing: state of the art, current trends and challenges (2022)Natural language processing: …Language Models as Knowledge Bases? (2019)Language Models as Knowledge …AutoML: A survey of the state-of-the-art (2020)AutoML: A survey of the state…Pre-trained models for natural language processing: A survey (2020)Pre-trained models for natura…
17 of 17 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

What cites it, inside the corpus

Topics

Topic ModelingComputer Science
Natural Language Processing TechniquesComputer Science
Speech Recognition and SynthesisComputer Science

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports6 author record(s) attached.
  • supports80 reference(s) recorded.
  • neutralThe DOI carries no year to check against.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:47+00:00.

sha256 a08467ae9504f237…