One voice is all you need: a one-shot approach to recognize your voice

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorNowshin, Priata
dc.contributor.authorDipto, Shahriar Rumi
dc.contributor.authorAhmed, Intesur
dc.contributor.authorChowdhury, Deboraj
dc.contributor.authorNoor, Galib Abdun.
dc.contributor.authorChakrabarty, Amitabha
dc.contributor.authorAbdullah M.T.
dc.contributor.authorRahman, Moshiur
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-07-27T03:44:06Z
dc.date.available2026-07-27T03:44:06Z
dc.date.issued2022-01-01
dc.description.abstractQuestion generation based on conversational context is a difficult problem to solve. A widely used technique for generating quality questions using fine-tuned models relies on a suitable answer and the context, usually the passage. But when it comes to conversational settings, the questions generated are not of the highest quality as they lack the contextual element in the question, especially due to the lack of co-reference resolution of the entity. Furthermore, in most of the evaluation techniques for generating questions, there seems to be a lack of utilizing powerful question-answering systems to judge the answerability of the questions generated. The most prevalent metric used for judging machine-generated text against the human gold standard, BLUE, unfortunately doesn't factor in whether a question answering system would be able to answer the question, but instead focuses mostly on the number of substrings that match against each other. Various question generation models following a generalized encoder-decoder architecture were evaluated using semantic textual similarity for both the generated questions and the generated answers. Although higher parameters in a model usually lend to better performance, our experiment displayed that such is not always the case, at least when there is a massive amount of context missing.
dc.description.versionPublished
dc.format.extent103-108
dc.identifier.citationP. Nowshin et al., "One Voice is All You Need: A One-Shot Approach to Recognize Your Voice," 2022 7th International Conference on Data Science and Machine Learning Applications (CDMA), Riyadh, Saudi Arabia, 2022, pp. 103-108, doi: 10.1109/CDMA54072.2022.00022.
dc.identifier.doi10.1109/CDMA54072.2022.00022
dc.identifier.issn9781665410144
dc.identifier.other2-s2.0-85127882657
dc.identifier.urihttps://hdl.handle.net/10361/28659
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/CDMA54072.2022.00022
dc.relation.ispartofProceedings 2022 7th International Conference on Data Science and Machine Learning Applications Cdma 2022
dc.relation.ispartofseriesProceedings 2022 7th International Conference on Data Science and Machine Learning Applications Cdma 2022
dc.relation.urihttps://ieeexplore.ieee.org/document/9736368
dc.subjectAudio classification
dc.subjectOne-shot learning
dc.subjectSiamese neural network
dc.subjectSpeaker recognition
dc.subjectTriplet loss
dc.subject.lcshAutomatic speech recognition.
dc.subject.lcshSignal processing--Digital techniques.
dc.titleOne voice is all you need: a one-shot approach to recognize your voice
dc.typeConference Proceeding
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameUniversity of Dhaka
person.affiliation.nameBRAC University
person.identifier.scopus-author-id57567460100
person.identifier.scopus-author-id59917694200
person.identifier.scopus-author-id57566860500
person.identifier.scopus-author-id57567270000
person.identifier.scopus-author-id57568261700
person.identifier.scopus-author-id35108854200
person.identifier.scopus-author-id57220902655
person.identifier.scopus-author-id58728070700

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
IMG_8345.jpg
Size:
27.35 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: