Abstract:Purpose: Malnutrition is a serious health concern, particularly among the older people living in residential aged care facilities. An automated and efficient method is required to identify the individuals afflicted with malnutrition in this setting. The recent advancements in transformer-based large language models (LLMs) equipped with sophisticated context-aware embeddings, such as RoBERTa, have significantly improved machine learning performance, particularly in predictive modelling. Enhancing the embeddings of these models on domain-specific corpora, such as clinical notes, is essential for elevating their performance in clinical tasks. Therefore, our study introduces a novel approach that trains a foundational RoBERTa model on nursing progress notes to develop a RAC domain-specific LLM. The model is further fine-tuned on nursing progress notes to enhance malnutrition identification and prediction in residential aged care setting. Methods: We develop our domain-specific model by training the RoBERTa LLM on 500,000 nursing progress notes from residential aged care electronic health records (EHRs). The model embeddings were used for two downstream tasks: malnutrition note identification and malnutrition prediction. Its performance was compared against baseline RoBERTa and BioClinicalBERT. Furthermore, we truncated long sequence text to fit into RoBERTa 512-token sequence length limitation, enabling our model to handle sequences up to1536 tokens. Results: Utilizing 5-fold cross-validation for both tasks, our RAC domain-specific LLM demonstrated significantly better performance over other models. In malnutrition note identification, it achieved a slightly higher F1-score of 0.966 compared to other LLMs. In prediction, it achieved significantly higher F1-score of 0.655. We enhanced our model predictive capability by integrating the risk factors extracted from each client notes, creating a combined data layer of structured risk factors and free-text notes. This integration improved the prediction performance, evidenced by an increased F1-score of 0.687. Conclusion: Our findings suggest that further fine-tuning a large language model on a domain-specific clinical corpus can improve the foundational model performance in clinical tasks. This specialized adaptation significantly improves our domain-specific model performance in tasks such as malnutrition risk identification and malnutrition prediction, making it useful for identifying and predicting malnutrition among older people living in residential aged care or long-term care facilities.

Leveraging Large Language Models for Metagenomic Analysis

ProkBERT family: genomic language models for microbiome applications

genomicBERT and data-free deep-learning model evaluation

Leveraging pre-trained language models for mining microbiome-disease relationships

Learning a deep language model for microbiomes: the power of large scale unlabeled microbiome data

Analyzing Large Microbiome Datasets Using Machine Learning and Big Data

FGBERT: Function-Driven Pre-trained Gene Language Model for Metagenomics

An Analysis on Large Language Models in Healthcare: A Case Study of BioBERT

Recent advances in deep learning and language models for studying the microbiome

Deciphering enzymatic potential in metagenomic reads through DNA language model

Advancing Plant Metabolic Research By Using Large Language Models To Expand Databases And Extract Labelled Data

Large-scale Machine Learning for Metagenomics Sequence Classification

Improving Taxonomic Classification with Feature Space Balancing

GeneGPT: augmenting large language models with domain tools for improved access to biomedical information

ProkBERT PhaStyle: Accurate Phage Lifestyle Prediction with Pretrained Genomic Language Models

ProtTrans: Towards Cracking the Language of Life's Code Through Self-Supervised Deep Learning and High Performance Computing

META$^\mathbf{2}$: Memory-efficient taxonomic classification and abundance estimation for metagenomics with deep learning

Fine-tuning large language models for effective nutrition support in residential aged care: a domain expertise approach

Using Large Language Models for Microbiome Findings Reports in Laboratory Diagnostics

Machine Learning and Deep Learning Applications in Metagenomic Taxonomy and Functional Annotation