Who Cited It

End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF

2016 · 2,595 citations · 2 from inside this corpus

Xuezhe Ma, Eduard Hovy

State-of-the-art sequence labeling systems traditionally require large amounts of taskspecific knowledge in the form of handcrafted features and data pre-processing. In this paper, we introduce a novel neutral network architecture that benefits from both word-and character-level representations automatically, by using combination of bidirectional LSTM, CNN and CRF. Our system is truly end-to-end, requiring no feature engineering or data preprocessing, thus making it applicable to a wide range of sequence labeling tasks. We evaluate our system on two data sets for two sequence labeling tasks -Penn Treebank WSJ corpus for part-of-speech (POS) tagging and CoNLL 2003 corpus for named entity recognition (NER). We obtain state-of-the-art performance on both datasets -97.55% accuracy for POS tagging and 91.21% F1 for NER.

End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF (2016)End-to-end Sequence Labeling …Long Short-Term Memory (1997)Long Short-Term MemoryAdam: A Method for Stochastic Optimization (2014)Adam: A Method for Stochastic…[No title in the source record — DROPS (Schloss Dagstuhl – Leibniz Center for Informatics… (2015)[No title in the source recor…Dropout: a simple way to prevent neural networks from overfitting (2014)Dropout: a simple way to prev…Glove: Global Vectors for Word Representation (2014)Glove: Global Vectors for Wor…Distributed Representations of Words and Phrases and their Compositionality (2013)Distributed Representations o…Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data (2001)Conditional Random Fields: Pr…Understanding the difficulty of training deep feedforward neural networks (2010)Understanding the difficulty …Speech recognition with deep recurrent neural networks (2013)Speech recognition with deep …Learning long-term dependencies with gradient descent is difficult (1994)Learning long-term dependenci…Building a Large Annotated Corpus of English: The Penn Treebank (1993)Building a Large Annotated Co…On the Properties of Neural Machine Translation: Encoder–Decoder Approaches (2014)On the Properties of Neural M…Learning to Forget: Continual Prediction with LSTM (2000)Learning to Forget: Continual…ADADELTA: An Adaptive Learning Rate Method (2012)ADADELTA: An Adaptive Learnin…Natural Language Processing (almost) from Scratch (2011)Natural Language Processing (…Neural Architectures for Named Entity Recognition (2016)Neural Architectures for Name…Natural Language Processing (almost) from Scratch (2011)Natural Language Processing (…On the difficulty of training Recurrent Neural Networks (2012)On the difficulty of training…Bidirectional LSTM-CRF Models for Sequence Tagging (2015)Bidirectional LSTM-CRF Models…Learning to forget: continual prediction with LSTM (1999)Learning to forget: continual…A Fast and Accurate Dependency Parser using Neural Networks (2014)A Fast and Accurate Dependenc…Introduction to the CoNLL-2002 Shared Task: Language-Independent Named Entity Recognition (2002)Introduction to the CoNLL-200…Design challenges and misconceptions in named entity recognition (2009)An Empirical Exploration of Recurrent Network Architectures (2015)An Empirical Exploration of R…A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data (2005)A Framework for Learning Pred…Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Langu… (2022)Pre-train, Prompt, and Predic…A Survey on Deep Learning for Named Entity Recognition (2020)A Survey on Deep Learning for…
27 of 27 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

PaperYearCited
Long Short-Term Memory1997101,359
Adam: A Method for Stochastic Optimization201484,698
[No title in the source record — DROPS (Schloss Dagstuhl – Leibniz Center for Informatics…201550,396
Dropout: a simple way to prevent neural networks from overfitting201434,236
Glove: Global Vectors for Word Representation201434,067
Distributed Representations of Words and Phrases and their Compositionality201318,054
Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data200112,994
Understanding the difficulty of training deep feedforward neural networks201012,673
Speech recognition with deep recurrent neural networks20138,916
Learning long-term dependencies with gradient descent is difficult19948,610
Building a Large Annotated Corpus of English: The Penn Treebank19937,538
On the Properties of Neural Machine Translation: Encoder–Decoder Approaches20146,683
Learning to Forget: Continual Prediction with LSTM20005,538
ADADELTA: An Adaptive Learning Rate Method20125,532
Natural Language Processing (almost) from Scratch20115,174
Neural Architectures for Named Entity Recognition20164,482
Natural Language Processing (almost) from Scratch20113,991
On the difficulty of training Recurrent Neural Networks20123,801
Bidirectional LSTM-CRF Models for Sequence Tagging20153,287
Learning to forget: continual prediction with LSTM19992,427
A Fast and Accurate Dependency Parser using Neural Networks20141,884
Introduction to the CoNLL-2002 Shared Task: Language-Independent Named Entity Recognition20021,583
Design challenges and misconceptions in named entity recognition20091,491
An Empirical Exploration of Recurrent Network Architectures20151,406
A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data20051,370

What cites it, inside the corpus

Topics

Natural Language Processing TechniquesComputer Science
Topic ModelingComputer Science
Speech Recognition and SynthesisComputer Science

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports2 author record(s) attached.
  • supports68 reference(s) recorded.
  • neutralThe DOI carries no year to check against.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:49+00:00.

sha256 a2172c1bbe00c44b…