Blind assistance: Object detection with voice feedback

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorKabir, Mosarrat Shazia
dc.contributor.authorKarishma Naaz, Syeda
dc.contributor.authorKabir, Md. Tahmid
dc.contributor.authorHussain, Md. Shahriar
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-22T05:27:02Z
dc.date.available2026-09-22T05:27:02Z
dc.date.issued2023-01-01
dc.description.abstractApplications of interaction between humans and computers, such as emergency response systems, heavily depend on emotion recognition. In this experiment, we present an investigation into Speech Emotion Recognition (SER) specifically in the context of emergency calls using a meticulously curated dataset. This dataset comprises audio recordings from 18 speakers, each expressing four distinct emotions (angry, drunk, painful, and stressful) while reading predefined emergency scenarios. We conducted comprehensive preprocessing, which involved feature extraction using Mel-Frequency Cepstral Coefficients (MFCCs), Chroma (Pitch Classes), and Mel Spectrogram Frequency. The performance of many machine learning models, such as the KNeighbors Classifier, MLP Classifier, Random Forest Classifier, Gradient Boost Classifier, SVM, and Logistic Regression, was then assessed by dividing the dataset into training and testing sets. Our results reveal that the KNeighbors Classifier outperforms other models, achieving an accuracy of 67.06% and maintaining balanced performance metrics. These results offer insightful information about whether SER is practical for emergency call applications. This research contributes to the understanding of emotion recognition in critical situations and can enhance the efficiency of emergency response systems by automating emotion assessment in distress calls. Our findings have practical implications for the development of intelligent systems that can better assist emergency service providers.
dc.description.versionPublished
dc.format.extent5 Pages
dc.identifier.citationM. S. Kabir, S. Karishma Naaz, M. T. Kabir and M. S. Hussain, "Blind Assistance: Object Detection with Voice Feedback," 2023 26th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2023, pp. 1-5, doi: 10.1109/ICCIT60459.2023.10440977.
dc.identifier.doi10.1109/ICCIT60459.2023.10440977
dc.identifier.issn9798350359015
dc.identifier.other2-s2.0-85187340986
dc.identifier.urihttps://hdl.handle.net/10361/30134
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/ICCIT60459.2023.10440977
dc.relation.ispartof2023 26th International Conference on Computer and Information Technology Iccit 2023
dc.relation.ispartofseries2023 26th International Conference on Computer and Information Technology Iccit 2023
dc.relation.urihttps://ieeexplore.ieee.org/document/10440977
dc.subjectEmergency call
dc.subjectEmotion classification
dc.subjectFeature extraction
dc.subjectSpeech emotion recognition
dc.subject.lcshPeople with visual disabilities--Means of communication--Technological innovations.
dc.subject.lcshImage processing--Digital techniques.
dc.titleBlind assistance: Object detection with voice feedback
dc.typeConference Proceeding
person.affiliation.nameNorth South University
person.affiliation.nameNorth South University
person.affiliation.nameBRAC University
person.affiliation.nameNorth South University
person.identifier.scopus-author-id58931067800
person.identifier.scopus-author-id58930875700
person.identifier.scopus-author-id58931452200
person.identifier.scopus-author-id55421370400

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
IMG_8345.jpg
Size:
27.35 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: