| International Journal of Computer Applications |
| Foundation of Computer Science (FCS), NY, USA |
| Volume 187 - Number 130 |
| Year of Publication: 2026 |
| Authors: Reddipalli Shashank, Gopi Kumar Jha, Nirbhay Singh, Ramesh K. Bhukya |
10.5120/ijca9045ccba6604
|
Reddipalli Shashank, Gopi Kumar Jha, Nirbhay Singh, Ramesh K. Bhukya . Psychoacoustic Framework for Automated Dysarthria Severity Assessment Via Hybrid Acoustic Ensembles. International Journal of Computer Applications. 187, 130 ( Jul 2026), 39-51. DOI=10.5120/ijca9045ccba6604
Dysarthria is a neurological motor speech disorder that affects speech production and intelligibility through impairments in articulation, phonation, and prosody. Reliable severity assessment is important for clinical diagnosis and long-term monitoring, yet existing automated systems rely heavily on Mel-Frequency Cepstral Coefficients (MFCCs), which may not adequately capture subtle pathological speech variations. This work investigates Bark-Frequency Cepstral Coefficients (BFCCs) as an alternative acoustic representation for dysarthria severity classification. A hybrid 193-dimensional feature vector is constructed by combining BFCCs with Mel-spectrogram, Chromagram, Spectral Contrast, and Tonnetz features to characterize complementary articulatory, spectral, and prosodic properties of dysarthric speech. Experiments are conducted on the TORGO corpus, with Voice Activity Detection (VAD) applied for silence removal and Adaptive Synthetic Sampling (ADASYN) used to address class imbalance. The extracted features are evaluated using a diverse set of machine learning (ML) models, including k-NN, Decision Tree (DT), Support Vector Machines (SVM), PCA, Random Forest (RF), AdaBoost, LogitBoost, CatBoost, LightGBM, XGBoost, and SGD. Results show that BFCC-based hybrid features provide strong discriminative capability across severity levels. Among the evaluated models, RF achieved the highest classification accuracy of 98.33%, followed by LightGBM (98.20%) and XGBoost (97.68%). The results indicate that BFCCs capture pathological speech characteristics more effectively than conventional cepstral representations and, when combined with ensemble learning, enable accurate and robust dysarthria severity classification.