Multitask audio analysis for emotion, gender, and speaker recognition in Bangla speech comparing features and models

Loading...
Thumbnail Image

Publisher

Institute of Electrical and Electronics Engineers Inc.

Citation

P. Deb, "Multitask Audio Analysis for Emotion, Gender, and Speaker Recognition in Bangla Speech Comparing Features and Models," 2023 International Conference on Advanced Computing Technologies and Applications (ICACTA), Mumbai, India, 2023, pp. 1-6, doi: 10.1109/ICACTA58201.2023.10393734.

Abstract

This study compares the effectiveness of various audio features and models for multitask audio analysis of Bangla speech, specifically novel approach for gender, and speaker recognition alongside emotion recognition. The study uses the SUBESCO dataset and extracts different features from the speech signals, including MFCC, Chroma, MEL, Rate of Zero-Crossings, and Spectral: Flux, Centroid, Roll-off. Different models, including Gradient Boosting Classifier, MLP Classifier, Random Forest Classifier, Logistic Regression, and SVM, are trained and tested for the three classification tasks. The results show that MFCC, Chroma, and MEL features are the most effective for multitask audio analysis of Bangla speech, achieving high accuracy for gender recognition, speaker recognition, and emotion recognition. Using the SUBESCO (Bangla) and RAVDESS (English) speech datasets, we conduct an evaluation of the suggested technique. The study provides valuable insights for future research in the field of speech analysis and suggests practical applications for speech recognition systems in various domains.

Description

Type

Conference Proceeding