Publicly Available Clinical
Emily Alsentzer, John R. Murphy, William Boag, Wei‐Hung Weng, Di Jindi low, Tristan Naumann, Matthew B. A. McDermott
Contextual word embedding models such as ELMo and BERT have dramatically improved performance for many natural language processing (NLP) tasks in recent months. However, these models have been minimally explored on specialty corpora, such as clinical text; moreover, in the clinical domain, no publicly-available pre-trained BERT models yet exist. In this work, we address this need by exploring and releasing BERT models for clinical text: one for generic clinical text and another for discharge summaries specifically. We demonstrate that using a domain-specific model yields performance improvements on 3/5 clinical NLP tasks, establishing a new state-of-the-art on the MedNLI dataset. We find that these domain-specific models are not as performant on 2 clinical de-identification tasks, and argue that this is a natural consequence of the differences between de-identified source text and synthetically non de-identified task text.
What this paper cites, inside the corpus
What cites it, inside the corpus
Links
Topics
| Topic Modeling | Computer Science |
| Natural Language Processing Techniques | Computer Science |
| Machine Learning in Healthcare | Computer Science |
Is this record sound?
partial
One field of this record is missing or disagrees with another. What is shown below is what the source publishes.
- supports7 author record(s) attached.
- supports20 reference(s) recorded.
- weakensThe DOI names 1909 but the record dates this to 2,019. One of the two is about a different paper.
- supportsA title is present.
Provenance
sha256 db1645b78a57e29a…