Rene: A Pre-trained Multi-modal Architecture for Auscultation of Respiratory Diseases

Pengfei Zhang,Zhihang Zheng,Shichen Zhang,Minghao Yang,Shaojun Tang
2024-06-07
Abstract:Compared with invasive examinations that require tissue sampling, respiratory sound testing is a non-invasive examination method that is safer and easier for patients to accept. In this study, we introduce Rene, a pioneering large-scale model tailored for respiratory sound recognition. Rene has been rigorously fine-tuned with an extensive dataset featuring a broad array of respiratory audio samples, targeting disease detection, sound pattern classification, and event identification. Our innovative approach applies a pre-trained speech recognition model to process respiratory sounds, augmented with patient medical records. The resulting multi-modal deep-learning framework addresses interpretability and real-time diagnostic challenges that have hindered previous respiratory-focused models. Benchmark comparisons reveal that Rene significantly outperforms existing models, achieving improvements of 10.27%, 16.15%, 15.29%, and 18.90% in respiratory event detection and audio classification on the SPRSound database. Disease prediction accuracy on the ICBHI database improved by 23% over the baseline in both mean average and harmonic scores. Moreover, we have developed a real-time respiratory sound discrimination system utilizing the Rene architecture. Employing state-of-the-art Edge AI technology, this system enables rapid and accurate responses for respiratory sound auscultation(<a class="link-external link-https" href="https://github.com/zpforlove/Rene" rel="external noopener nofollow">this https URL</a>).
Sound,Artificial Intelligence,Audio and Speech Processing,Quantitative Methods
What problem does this paper attempt to address?
The main problem that this paper attempts to solve is to develop a non - invasive and efficient method for analyzing lung auscultation sounds, in order to achieve more accurate detection and diagnosis of respiratory diseases. Specifically, the paper introduces a multimodal pre - training framework named RENE, which aims to improve the accuracy of disease detection, sound pattern classification and event recognition by processing breathing sounds. ### Main problems: 1. **Requirement for non - invasive examination**: Compared with invasive examinations that require tissue sampling, the breathing sound test is a non - invasive examination method that is safer and more acceptable to patients. 2. **Limitations of existing models**: Existing models for breathing sound analysis face challenges in interpretability and real - time diagnosis and cannot meet the needs of clinical applications. 3. **Multimodal information fusion**: Relying solely on a single model for disease diagnosis is not sufficient. Combining multiple information sources (such as breathing sounds and medical records) can improve the robustness and interpretability of diagnosis. ### Solutions: - **RENE architecture**: RENE is a large - scale pre - trained model, which is fine - tuned and specifically used for breathing sound recognition. It combines a speech recognition model and a patient's medical records to form a multimodal deep - learning framework. - **Dataset integration**: The paper uses two publicly available breathing sound databases (SPRSound and ICBHI) and integrates them into a training set containing 16,300 breathing sound recordings and a test set of 7,979 recordings. - **Feature extraction and fusion**: The RENE architecture includes two main modules: a clinical record module and a feature extraction module. The former extracts features from structured medical records, and the latter uses the Whisper model and the Conformer network to extract features from breathing sounds. Finally, by fusing the outputs of these two modules, the final diagnosis result is generated. - **Real - time system development**: Based on the RENE architecture, a real - time breathing sound discrimination system is developed, which uses edge AI technology to achieve a fast and accurate response. ### Achievements: - **Performance improvement**: RENE significantly outperforms existing models in multiple benchmark tests. For example, in the breathing event detection and audio classification tasks on the SPRSound database, it improves performance by 10.27%, 16.15%, 15.29% and 18.90% respectively. - **Disease prediction accuracy**: On the ICBHI database, the disease prediction accuracy of RENE is 23% higher than that of the baseline model, especially outstanding in the average score and the harmonic average score. In conclusion, this paper solves the challenges of interpretability and real - time diagnosis in breathing sound analysis by introducing the RENE architecture and shows its potential in clinical applications.