Who Cited It

Understanding the difficulty of training deep feedforward neural networks

2010 · 12,673 citations · 27 from inside this corpus

Xavier Glorot, Yoshua Bengio

The source holds an abstract for this work, but its best open-access copy is under no open licence, which does not permit us to republish the text. Read it at the source below.

Understanding the difficulty of training deep feedforward neural networks (2010)Understanding the difficulty …Learning representations by back-propagating errors (1986)Learning representations by b…A Fast Learning Algorithm for Deep Belief Nets (2006)A Fast Learning Algorithm for…Learning long-term dependencies with gradient descent is difficult (1994)Learning long-term dependenci…A unified architecture for natural language processing (2008)A unified architecture for na…Learning Deep Architectures for AI (2009)Learning Deep Architectures f…Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Sh… (2015)Batch Normalization: Accelera…Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Fou… (2012)Deep Neural Networks for Acou…Deep Sparse Rectifier Neural Networks (2011)Deep Sparse Rectifier Neural …Translating Embeddings for Modeling Multi-relational Data (2013)Translating Embeddings for Mo…Deep Learning in Medical Image Analysis (2017)Deep Learning in Medical Imag…Learning without Forgetting (2017)Learning without ForgettingDeep Convolutional Neural Networks for Image Classification: A Comprehensive Review (2017)Deep Convolutional Neural Net…On the importance of initialization and momentum in deep learning (2013)On the importance of initiali…Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition (2011)Context-Dependent Pre-Trained…A systematic study of the class imbalance problem in convolutional neural networks (2018)A systematic study of the cla…Deeper Insights Into Graph Convolutional Networks for Semi-Supervised Learning (2018)Deeper Insights Into Graph Co…End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF (2016)End-to-end Sequence Labeling …Barren plateaus in quantum neural network training landscapes (2018)Barren plateaus in quantum ne…On the Opportunities and Risks of Foundation Models (2021)On the Opportunities and Risk…Ensemble deep learning: A review (2022)Ensemble deep learning: A rev…Practical Recommendations for Gradient-Based Training of Deep Architectures (2012)Deep Neural Networks for Acoustic Modeling in Speech Recognition (2012)Deep Neural Networks for Acou…Review of Deep Learning Algorithms and Architectures (2019)Review of Deep Learning Algor…Deep Learning: Methods and Applications (2014)Deep Learning: Methods and Ap…A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer … (2017)A Gift from Knowledge Distill…Convolutional, Long Short-Term Memory, fully connected Deep Neural Networks (2015)Convolutional, Long Short-Ter…A State-of-the-Art Survey on Deep Learning Theory and Architectures (2019)A State-of-the-Art Survey on …Deep Reinforcement Learning That Matters (2018)Deep Reinforcement Learning T…Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach (2017)Making Deep Neural Networks R…Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activa… (2016)Quantized Neural Networks: Tr…Learning Temporal Regularity in Video Sequences (2016)Learning Temporal Regularity …A Survey on Deep Learning (2018)A Survey on Deep Learning
32 of 32 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

What cites it, inside the corpus

PaperYearCited
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Sh…201524,404
Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Fou…201210,399
Deep Sparse Rectifier Neural Networks20115,428
Translating Embeddings for Modeling Multi-relational Data20135,188
Deep Learning in Medical Image Analysis20174,917
Learning without Forgetting20174,034
Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review20173,570
On the importance of initialization and momentum in deep learning20133,523
Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition20113,089
A systematic study of the class imbalance problem in convolutional neural networks20183,072
Deeper Insights Into Graph Convolutional Networks for Semi-Supervised Learning20182,652
End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF20162,595
Barren plateaus in quantum neural network training landscapes20182,289
On the Opportunities and Risks of Foundation Models20212,265
Ensemble deep learning: A review20222,136
Practical Recommendations for Gradient-Based Training of Deep Architectures20121,960
Deep Neural Networks for Acoustic Modeling in Speech Recognition20121,907
Review of Deep Learning Algorithms and Architectures20191,878
Deep Learning: Methods and Applications20141,801
A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer …20171,660
Convolutional, Long Short-Term Memory, fully connected Deep Neural Networks20151,646
A State-of-the-Art Survey on Deep Learning Theory and Architectures20191,619
Deep Reinforcement Learning That Matters20181,580
Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach20171,464
Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activa…20161,420
Learning Temporal Regularity in Video Sequences20161,407
A Survey on Deep Learning20181,398

Topics

Neural Networks and ApplicationsComputer Science
Generative Adversarial Networks and Image SynthesisComputer Science
Gaussian Processes and Bayesian InferenceComputer Science

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports2 author record(s) attached.
  • supports17 reference(s) recorded.
  • neutralThe DOI carries no year to check against.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:40+00:00.

sha256 7e3d99a592f7f61f…