Vision meets language: Multimodal transformers elevating predictive power in visual question answering

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorKhandaker, Sajidul Islam
dc.contributor.authorTalukdar, Tahmina
dc.contributor.authorSarker, Prima
dc.contributor.authorMehedi, Md Humaion Kabir
dc.contributor.authorRhythm, Ehsanur Rahman
dc.contributor.authorRasel, Annajiat Alim
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-24T10:23:45Z
dc.date.available2026-09-24T10:23:45Z
dc.date.issued2023-01-01
dc.description.abstractVisual Question Answering (VQA) is a field where computer vision and natural language processing intersect to develop systems capable of comprehending visual information and answering natural language questions. In visual question answering , algorithms interpret real-world images in response to questions expressed in human language. Our paper presents an extensive experimental study on Visual Question Answering (VQA) using a diverse set of multimodal transformers. The VQA task requires systems to comprehend both visual content and natural language questions. To address this challenge, we explore the performance of various pre-trained transformer architectures for encoding questions, including BERT, RoBERTa, and ALBERT, as well as image transformers, such as ViT, DeiT, and BEiT, for encoding images. Multimodal transformers' smooth fusion of visual and text data promotes cross-modal understanding and strengthens reasoning skills. On benchmark datasets like the Visual Question Answering (VQA) v2.0 dataset, we rigorously test and fine-tune these models to assess their effectiveness and compare their performance to more conventional VQA methods. The results show that multimodal transformers significantly outperform traditional techniques in terms of performance. Additionally, the models' attention maps give users insights into how they make decisions, improving interpretability and comprehension. Because of their adaptability, the tested transformer topologies have the potential to be used in a wide range of VQA applications, such as robotics, healthcare, and assistive technology. This study demonstrates the effectiveness and promise of multimodal transformers as a method for improving the effectiveness of visual question-answering systems.
dc.description.versionPublished
dc.format.extent8 Pages
dc.identifier.citationS. I. Khandaker, T. Talukdar, P. Sarker, M. H. K. Mehedi, E. R. Rhythm and A. A. Rasel, "Vision Meets Language: Multimodal Transformers Elevating Predictive Power in Visual Question Answering," 2023 26th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2023, pp. 1-6, doi: 10.1109/ICCIT60459.2023.10441514.
dc.identifier.doi10.1109/ICCIT60459.2023.10441514
dc.identifier.issn9798350359015
dc.identifier.other2-s2.0-85187408076
dc.identifier.urihttps://hdl.handle.net/10361/30226
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/ICCIT60459.2023.10441514
dc.relation.ispartof2023 26th International Conference on Computer and Information Technology Iccit 2023
dc.relation.ispartofseries2023 26th International Conference on Computer and Information Technology Iccit 2023
dc.relation.urihttps://ieeexplore.ieee.org/document/10441514
dc.subjectVisualization
dc.subjectImage coding
dc.subjectComputational modeling
dc.subjectMedical services
dc.subjectPredictive models
dc.subjectTransformers
dc.subjectQuestion answering (information retrieval)
dc.subjectVisual Question Answering (VQA)
dc.subjectBenchmark datasets
dc.subjectMultimodal transformers
dc.subjectInterpretability
dc.subject.lcshNatural language processing (Computer science).
dc.titleVision meets language: Multimodal transformers elevating predictive power in visual question answering
dc.typeConference Proceeding
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.identifier.scopus-author-id58930093300
person.identifier.scopus-author-id58930093400
person.identifier.scopus-author-id58931446600
person.identifier.scopus-author-id57422283000
person.identifier.scopus-author-id57971901600
person.identifier.scopus-author-id56495276900

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Demo.pdf.jpg
Size:
2.36 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: