Abstract:Aphasia is a type of speech disorder that can cause speech defects in a person. Identifying the severity level of the aphasia patient is critical for the rehabilitation process. In this research, we identify ten aphasia severity levels motivated by specific speech therapies based on the presence or absence of identified characteristics in aphasic speech in order to give more specific treatment to the patient. In the aphasia severity level classification process, we experiment on different speech feature extraction techniques, lengths of input audio samples, and machine learning classifiers toward classification performance. Aphasic speech is required to be sensed by an audio sensor and then recorded and divided into audio frames and passed through an audio feature extractor before feeding into the machine learning classifier. According to the results, the mel frequency cepstral coefficient (MFCC) is the most suitable audio feature extraction method for the aphasic speech level classification process, as it outperformed the classification performance of all mel-spectrogram, chroma, and zero crossing rates by a large margin. Furthermore, the classification performance is higher when 20 s audio samples are used compared with 10 s chunks, even though the performance gap is narrow. Finally, the deep neural network approach resulted in the best classification performance, which was slightly better than both K-nearest neighbor (KNN) and random forest classifiers, and it was significantly better than decision tree algorithms. Therefore, the study shows that aphasia level classification can be completed with accuracy, precision, recall, and F1-score values of 0.99 using MFCC for 20 s audio samples using the deep neural network approach in order to recommend corresponding speech therapy for the identified level. A web application was developed for English-speaking aphasia patients to self-diagnose the severity level and engage in speech therapies.

Aphasic Speech Recognition using a Mixture of Speech Intelligibility Experts

Quality-aware Aggregated Conformal Prediction for Silent Speech Recognition

An optimal hybrid AI-ResNet for accurate severity detection and classification of patients with aphasia disorder

Use of Speech Impairment Severity for Dysarthric Speech Recognition

A Strategic Approach for Robust Dysarthric Speech Recognition

Speech Enhancement using a Deep Mixture of Experts

Careful Whisper -- leveraging advances in automatic speech recognition for robust and interpretable aphasia subtype classification

Speech Emotion Recognition Based on Syllable-Level Feature Extraction

Real-time Speech Emotion Recognition Based on Syllable-Level Feature Extraction

Speech Enhancement Based on Deep Mixture of Distinguishing Experts

SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR

Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model

An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement

Automatic Assessment of Aphasic Speech Sensed by Audio Sensors for Classification into Aphasia Severity Levels to Recommend Speech Therapies

Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech

Integrating Articulatory Features into Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English Speech.

Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models

MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction

Latent Phrase Matching for Dysarthric Speech

Automatic recognition and detection of aphasic natural speech

A New Benchmark of Aphasia Speech Recognition and Detection Based on E-Branchformer and Multi-task Learning