Abstract:Deep learning techniques have shown promising results in the automatic classification of respiratory sounds. However, accurately distinguishing these sounds in real-world noisy conditions poses challenges for clinical deployment. Additionally, predicting signals with only background noise could undermine user trust in the system. This paper aims to investigate the feasibility and effectiveness of incorporating a deep learning-based audio enhancement preprocessing step into automatic respiratory sound classification systems to improve robustness and clinical applicability. Multiple experiments were conducted using different audio enhancement model structures and classification models. The classification performance was compared to the baseline method of noise injection data augmentation. Experiments were performed on two datasets: the ICBHI respiratory sound dataset, which includes 5.5 hours of recordings, and the Formosa Archive of Breath Sounds (FABS) dataset, comprising 14.6 hours of recordings. Additionally, a physician validation study was conducted by 7 senior physicians to assess the clinical utility of the <a class="link-external link-http" href="http://system.The" rel="external noopener nofollow">this http URL</a> integration of the audio enhancement pipeline resulted in a 21.88% increase in the ICBHI classification score on the ICBHI dataset and a 4.10% improvement on the FABS dataset in multi-class noisy scenarios. Quantitative analysis from the physician validation study revealed improvements in efficiency, diagnostic confidence, and trust during model-assisted diagnosis, with workflows integrating enhanced audio leading to an 11.61% increase in diagnostic sensitivity and facilitating high-confidence diagnoses. Incorporating an audio enhancement algorithm significantly enhances the robustness and clinical utility of automatic respiratory sound classification systems, improving performance in noisy environments and fostering greater trust among medical professionals.

Patch-level Contrastive Embedding Learning for Respiratory Sound Classification

Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification

Automatic Respiratory Sound Classification Via Multi-Branch Temporal Convolutional Network

Stethoscope-guided Supervised Contrastive Learning for Cross-domain Adaptation on Respiratory Sound Classification

A Contrastive Embedding-Based Domain Adaptation Method for Lung Sound Recognition in Children Community-Acquired Pneumonia

SoundCLR: Contrastive Learning of Representations For Improved Environmental Sound Classification

Pretraining Respiratory Sound Representations using Metadata and Contrastive Learning

Joint Energy-based Model for Semi-supervised Respiratory Sound Classification: A Method of Insensitive to Distribution Mismatch

Adversarial Fine-tuning using Generated Respiratory Sound to Address Class Imbalance

Improving Robustness and Clinical Applicability of Automatic Respiratory Sound Classification Using Deep Learning-Based Audio Enhancement: Algorithm Development and Validation Study

Supervised Contrastive Learning Framework and Hardware Implementation of Learned ResNet for Real-time Respiratory Sound Classification

Classify Respiratory Abnormality in Lung Sounds Using STFT and a Fine-Tuned ResNet18 Network

An Inception-Residual-Based Architecture with Multi-Objective Loss for Detecting Respiratory Anomalies

CNN-MoE based framework for classification of respiratory anomalies and lung disease detection

Deep Neural Network for Respiratory Sound Classification in Wearable Devices Enabled by Patient Specific Model Tuning

Exploring Self-Supervised Contrastive Learning of Spatial Sound Event Representation

Mask Detection and Breath Monitoring from Speech: on Data Augmentation, Feature Representation and Modeling

RepAugment: Input-Agnostic Representation-Level Augmentation for Respiratory Sound Classification

Intelligent Stethoscope Using Full Self-Attention Mechanism for Abnormal Respiratory Sound Recognition

Curricular Contrastive Regularization for Speech Enhancement with Self-Supervised Representations.

Fully Few-shot Class-incremental Audio Classification Using Expandable Dual-embedding Extractor