Abstract:Motivation: The explosive increase of biomedical literature has made information extraction an increasingly important tool for biomedical research. A fundamental task is the recognition of biomedical named entities in text (BNER) such as genes/proteins, diseases and species. Recently, a domain-independent method based on deep learning and statistical word embeddings, called long short-term memory network-conditional random field (LSTM-CRF), has been shown to outperform state-of-the-art entity-specific BNER tools. However, this method is dependent on gold-standard corpora (GSCs) consisting of hand-labeled entities, which tend to be small but highly reliable. An alternative to GSCs are silver-standard corpora (SSCs), which are generated by harmonizing the annotations made by several automatic annotation systems. SSCs typically contain more noise than GSCs but have the advantage of containing many more training examples. Ideally, these corpora could be combined to achieve the benefits of both, which is an opportunity for transfer learning. In this work, we analyze to what extent transfer learning improves upon state-of-the-art results for BNER.Results: We demonstrate that transferring a deep neural network (DNN) trained on a large, noisy SSC to a smaller, but more reliable GSC significantly improves upon state-of-the-art results for BNER. Compared to a state-of-the-art baseline evaluated on 23 GSCs covering four different entity classes, transfer learning results in an average reduction in error of approximately 11%. We found transfer learning to be especially beneficial for target datasets with a small number of labels (approximately 6000 or less).Availability and implementation: Source code for the LSTM-CRF is available at https://github.com/Franck-Dernoncourt/NeuroNER/ and links to the corpora are available at https://github.com/BaderLab/Transfer-Learning-BNER-Bioinformatics-2018/.Supplementary information: Supplementary data are available at Bioinformatics online.

Transfer Learning for Low-Resource Clinical Named Entity Recognition

Enhancing Low Resource NER Using Assisting Language And Transfer Learning

Language inference-based learning for Low-Resource Chinese clinical named entity recognition using language model

Low-Resource Named Entity Recognition with Cross-Lingual, Character-Level Neural Conditional Random Fields

Dynamic Transfer Learning for Named Entity Recognition

Ensemble Transfer Learning on Augmented Domain Resources for Oncological Named Entity Recognition in Chinese Clinical Records

Neural Machine Translation of Clinical Text: An Empirical Investigation into Multilingual Pre-Trained Language Models and Transfer-Learning

Data augmentation and transfer learning for cross-lingual Named Entity Recognition in the biomedical domain

Converse Attention Knowledge Transfer for Low-Resource Named Entity Recognition

Analysing Cross-Lingual Transfer in Low-Resourced African Named Entity Recognition

Transfer learning for biomedical named entity recognition with neural networks

A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity Recognizers

Cross-Lingual Transfer for Distantly Supervised and Low-resources Indonesian NER

Effective Transfer Learning for Low-Resource Natural Language Understanding

Low-Resource Adaptation of Neural NLP Models

Med7: a transferable clinical natural language processing model for electronic health records

Study on Application of Transfer Learning in Entity Recognition of Low Resource Environment

Neural Cross-Lingual Named Entity Recognition with Minimal Resources

Accelerating Clinical Text Annotation in Underrepresented Languages: A Case Study on Text De-Identification

Embedding Transfer for Low-Resource Medical Named Entity Recognition: A Case Study on Patient Mobility

Few-shot clinical entity recognition in English, French and Spanish: masked language models outperform generative model prompting