Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov low, Guoguo Chen low, Daniel Povey, Sanjeev Khudanpur
This paper introduces a new corpus of read English speech, suitable for training and evaluating speech recognition systems. The LibriSpeech corpus is derived from audiobooks that are part of the LibriVox project, and contains 1000 hours of speech sampled at 16 kHz. We have made the corpus freely available for download, along with separately prepared language-model training data and pre-built language models. We show that acoustic models trained on LibriSpeech give lower error rate on the Wall Street Journal (WSJ) test sets than models trained on WSJ itself. We are also releasing Kaldi scripts that make it easy to build these systems.
What this paper cites, inside the corpus
| Paper | Year | Cited |
|---|---|---|
| Identification of common molecular subsequences | 1981 | 10,144 |
| Kaldi Speech Recognition Toolkit | 2024 | 4,895 |
| SRILM - an extensible language modeling toolkit | 2002 | 4,404 |
| Improved backing-off for M-gram language modeling | 2002 | 1,500 |
What cites it, inside the corpus
Links
Topics
| Speech Recognition and Synthesis | Computer Science |
| Music and Audio Processing | Computer Science |
| Speech and Audio Processing | Computer Science |
Is this record sound?
complete
Nothing in this record contradicts itself and no field we check is missing.
- supports4 author record(s) attached.
- supports29 reference(s) recorded.
- supportsThe DOI's year agrees with the publication year.
- supportsA title is present.
Provenance
sha256 5cad55ac5d4d41e1…