Who Cited It

ADADELTA: An Adaptive Learning Rate Method

2012 · arXiv (Cornell University) · 5,532 citations · 24 from inside this corpus

Matthew D. Zeiler low

The source holds an abstract for this work, but its best open-access copy is under no open licence, which does not permit us to republish the text. Read it at the source below.

ADADELTA: An Adaptive Learning Rate Method (2012)ADADELTA: An Adaptive Learnin…Learning representations by back-propagating errors (1986)Learning representations by b…Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition (2011)Context-Dependent Pre-Trained…Adam: A Method for Stochastic Optimization (2014)Adam: A Method for Stochastic…Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Transla… (2014)Learning Phrase Representatio…Deep learning in neural networks: An overview (2014)Deep learning in neural netwo…Neural Machine Translation by Jointly Learning to Align and Translate (2014)Neural Machine Translation by…Convolutional Neural Networks for Sentence Classification (2014)Convolutional Neural Networks…Domain-Adversarial Training of Neural Networks (2017)Domain-Adversarial Training o…[No title in the source record — Edinburgh Research Explorer (University of Edinburgh)][No title in the source recor…On the Properties of Neural Machine Translation: Encoder–Decoder Approaches (2014)On the Properties of Neural M…An overview of gradient descent optimization algorithms (2016)An overview of gradient desce…Neural Architectures for Named Entity Recognition (2016)Neural Architectures for Name…Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review (2017)Deep Convolutional Neural Net…Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Transla… (2014)Learning Phrase Representatio…End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF (2016)End-to-end Sequence Labeling …Optimization as a Model for Few-Shot Learning (2017)Optimization as a Model for F…Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond (2016)Abstractive Text Summarizatio…Attention-Based Bidirectional Long Short-Term Memory Networks for Relation Classification (2016)Attention-Based Bidirectional…Knowledge Graph Embedding via Dynamic Mapping Matrix (2015)Knowledge Graph Embedding via…Deep Contextualized Word Representations (2018)Deep Contextualized Word Repr…SGDR: Stochastic Gradient Descent with Warm Restarts (2016)SGDR: Stochastic Gradient Des…Asynchronous Methods for Deep Reinforcement Learning (2016)Asynchronous Methods for Deep…A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer … (2017)A Gift from Knowledge Distill…Automated Machine Learning (2019)Automated Machine LearningImproved Adam Optimizer for Deep Neural Networks (2018)Improved Adam Optimizer for D…A Survey on Deep Learning (2018)A Survey on Deep Learning
26 of 26 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

What cites it, inside the corpus

PaperYearCited
Adam: A Method for Stochastic Optimization201484,698
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Transla…201425,048
Deep learning in neural networks: An overview201418,236
Neural Machine Translation by Jointly Learning to Align and Translate201414,615
Convolutional Neural Networks for Sentence Classification201413,987
Domain-Adversarial Training of Neural Networks20177,702
[No title in the source record — Edinburgh Research Explorer (University of Edinburgh)]7,291
On the Properties of Neural Machine Translation: Encoder–Decoder Approaches20146,683
An overview of gradient descent optimization algorithms20164,807
Neural Architectures for Named Entity Recognition20164,482
Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review20173,570
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Transla…20143,518
End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF20162,595
Optimization as a Model for Few-Shot Learning20172,436
Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond20162,233
Attention-Based Bidirectional Long Short-Term Memory Networks for Relation Classification20162,062
Knowledge Graph Embedding via Dynamic Mapping Matrix20151,868
Deep Contextualized Word Representations20181,783
SGDR: Stochastic Gradient Descent with Warm Restarts20161,735
Asynchronous Methods for Deep Reinforcement Learning20161,689
A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer …20171,660
Automated Machine Learning20191,451
Improved Adam Optimizer for Deep Neural Networks20181,408
A Survey on Deep Learning20181,398

Topics

Neural Networks and ApplicationsComputer Science
Domain Adaptation and Few-Shot LearningComputer Science
Machine Learning and AlgorithmsComputer Science

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports1 author record(s) attached.
  • supports6 reference(s) recorded.
  • neutralThe DOI carries no year to check against.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:44+00:00.

sha256 648f8491aaceaa34…