Multitask audio analysis for emotion, gender, and speaker recognition in Bangla speech comparing features and models
| bracu.type.group | Research Publications | |
| datacite.rights | Metadata Only | |
| dc.contributor.author | Deb, Priom | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-09-01T08:03:26Z | |
| dc.date.available | 2026-09-01T08:03:26Z | |
| dc.date.issued | 2023-01-01 | |
| dc.description.abstract | This study compares the effectiveness of various audio features and models for multitask audio analysis of Bangla speech, specifically novel approach for gender, and speaker recognition alongside emotion recognition. The study uses the SUBESCO dataset and extracts different features from the speech signals, including MFCC, Chroma, MEL, Rate of Zero-Crossings, and Spectral: Flux, Centroid, Roll-off. Different models, including Gradient Boosting Classifier, MLP Classifier, Random Forest Classifier, Logistic Regression, and SVM, are trained and tested for the three classification tasks. The results show that MFCC, Chroma, and MEL features are the most effective for multitask audio analysis of Bangla speech, achieving high accuracy for gender recognition, speaker recognition, and emotion recognition. Using the SUBESCO (Bangla) and RAVDESS (English) speech datasets, we conduct an evaluation of the suggested technique. The study provides valuable insights for future research in the field of speech analysis and suggests practical applications for speech recognition systems in various domains. | |
| dc.description.version | Published | |
| dc.format.extent | 6 Pages | |
| dc.identifier.citation | P. Deb, "Multitask Audio Analysis for Emotion, Gender, and Speaker Recognition in Bangla Speech Comparing Features and Models," 2023 International Conference on Advanced Computing Technologies and Applications (ICACTA), Mumbai, India, 2023, pp. 1-6, doi: 10.1109/ICACTA58201.2023.10393734. | |
| dc.identifier.doi | 10.1109/ICACTA58201.2023.10393734 | |
| dc.identifier.issn | 9798350348347 | |
| dc.identifier.other | 2-s2.0-85184810321 | |
| dc.identifier.uri | https://hdl.handle.net/10361/29657 | |
| dc.language.iso | en_US | |
| dc.publisher | Institute of Electrical and Electronics Engineers Inc. | |
| dc.relation.hasversion | 10.1109/ICACTA58201.2023.10393734 | |
| dc.relation.ispartof | Proceedings of 3rd International Conference on Advanced Computing Technologies and Applications Icacta 2023 | |
| dc.relation.ispartofseries | Proceedings of 3rd International Conference on Advanced Computing Technologies and Applications Icacta 2023 | |
| dc.relation.uri | https://ieeexplore.ieee.org/document/10393734 | |
| dc.subject | Support vector machines | |
| dc.subject | Emotion recognition | |
| dc.subject | Analytical models | |
| dc.subject | Speech analysis | |
| dc.subject | Speech recognition | |
| dc.subject | Feature extraction | |
| dc.subject | Speaker recognition | |
| dc.subject | Bangla speech | |
| dc.subject | Audio features | |
| dc.subject | Gender recognition | |
| dc.subject | Speaker recognition | |
| dc.subject | Emotion recognition | |
| dc.subject.lcsh | Speech processing systems. | |
| dc.subject.lcsh | Automatic speech recognition. | |
| dc.title | Multitask audio analysis for emotion, gender, and speaker recognition in Bangla speech comparing features and models | |
| dc.type | Conference Proceeding | |
| person.affiliation.name | BRAC University | |
| person.identifier.scopus-author-id | 58882099300 |