Multitask audio analysis for emotion, gender, and speaker recognition in Bangla speech comparing features and models

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorDeb, Priom
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-01T08:03:26Z
dc.date.available2026-09-01T08:03:26Z
dc.date.issued2023-01-01
dc.description.abstractThis study compares the effectiveness of various audio features and models for multitask audio analysis of Bangla speech, specifically novel approach for gender, and speaker recognition alongside emotion recognition. The study uses the SUBESCO dataset and extracts different features from the speech signals, including MFCC, Chroma, MEL, Rate of Zero-Crossings, and Spectral: Flux, Centroid, Roll-off. Different models, including Gradient Boosting Classifier, MLP Classifier, Random Forest Classifier, Logistic Regression, and SVM, are trained and tested for the three classification tasks. The results show that MFCC, Chroma, and MEL features are the most effective for multitask audio analysis of Bangla speech, achieving high accuracy for gender recognition, speaker recognition, and emotion recognition. Using the SUBESCO (Bangla) and RAVDESS (English) speech datasets, we conduct an evaluation of the suggested technique. The study provides valuable insights for future research in the field of speech analysis and suggests practical applications for speech recognition systems in various domains.
dc.description.versionPublished
dc.format.extent6 Pages
dc.identifier.citationP. Deb, "Multitask Audio Analysis for Emotion, Gender, and Speaker Recognition in Bangla Speech Comparing Features and Models," 2023 International Conference on Advanced Computing Technologies and Applications (ICACTA), Mumbai, India, 2023, pp. 1-6, doi: 10.1109/ICACTA58201.2023.10393734.
dc.identifier.doi10.1109/ICACTA58201.2023.10393734
dc.identifier.issn9798350348347
dc.identifier.other2-s2.0-85184810321
dc.identifier.urihttps://hdl.handle.net/10361/29657
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/ICACTA58201.2023.10393734
dc.relation.ispartofProceedings of 3rd International Conference on Advanced Computing Technologies and Applications Icacta 2023
dc.relation.ispartofseriesProceedings of 3rd International Conference on Advanced Computing Technologies and Applications Icacta 2023
dc.relation.urihttps://ieeexplore.ieee.org/document/10393734
dc.subjectSupport vector machines
dc.subjectEmotion recognition
dc.subjectAnalytical models
dc.subjectSpeech analysis
dc.subjectSpeech recognition
dc.subjectFeature extraction
dc.subjectSpeaker recognition
dc.subjectBangla speech
dc.subjectAudio features
dc.subjectGender recognition
dc.subjectSpeaker recognition
dc.subjectEmotion recognition
dc.subject.lcshSpeech processing systems.
dc.subject.lcshAutomatic speech recognition.
dc.titleMultitask audio analysis for emotion, gender, and speaker recognition in Bangla speech comparing features and models
dc.typeConference Proceeding
person.affiliation.nameBRAC University
person.identifier.scopus-author-id58882099300

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
IMG_8345.jpg
Size:
27.35 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: