Lightweight Visual Question Answering (VQA) model for skin disease detection

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorNoor, Abtahi
dc.contributor.authorMahe, Raiyan Habib
dc.contributor.authorAziz, Azwad
dc.contributor.authorChakrabarty, Amitabha
dc.contributor.authorTasin, Ridwan Noor
dc.contributor.authorIslam, Md Fahim Ul
dc.contributor.authorRahman, Rafeed
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-08-22T04:14:21Z
dc.date.available2026-08-22T04:14:21Z
dc.date.issued2025-01-01
dc.description.abstractVisual Question Answering (VQA) is an area of artificial intelligence that combines image analysis with natural language understanding to generate context-aware answers to queries. Its application in the medical field, particularly in dermatology, holds significant importance by enabling accessible, interpretable, and efficient diagnostic support. Currently, there are plenty of good skin disease classification datasets, however, there is a lack of structured VQA datasets for skin diseases that can be used for training as well as benchmarking models. Hence, we developed a custom dataset of 1,038 images for 11 disease classes, with seven question-answer pairs per image. Moreover, existing VQA models are extremely heavy weight and require specialized hardware to train and run. To address this, our research proposes a lightweight VQA model pipeline capable of identifying common skin diseases from images and responding to clinically relevant questions related to disease name, severity, causes, diagnostic approach, prevention, contagiousness, and cancer risk. The model uses a modular architecture that integrates a Vision Transformer (ViT) with 86 million parameters for image encoding and MiniLM, a transformer-based text encoder with 22 million parameters. It achieved a high accuracy of 94.87% while minimizing computational requirements. In addition, we have also trained state-of-the-art vision language models such as Gemma-3, QwenVl-2.5, LLaVA-1.5, and BLIP-2 using our dataset for comparison. Among these, BLIP-2 achieved the highest Sentence-BERT score of 81.41%, indicating strong alignment between predicted and reference answers.
dc.description.versionPublished
dc.format.extent6 Pages
dc.identifier.citationA. Noor et al., "Lightweight Visual Question Answering (VQA) Model for Skin Disease Detection," 2025 7th International Conference on Electrical Information and Communication Technology (EICT), Khulna, Bangladesh, 2025, pp. 1-6, doi: 10.1109/EICT68394.2025.11355641.
dc.identifier.doi10.1109/EICT68394.2025.11355641
dc.identifier.issn9798331593926
dc.identifier.other2-s2.0-105033533174
dc.identifier.urihttps://hdl.handle.net/10361/29411
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/EICT68394.2025.11355641
dc.relation.ispartof2025 7th International Conference on Electrical Information and Communication Technology Eict 2025
dc.relation.ispartofseries2025 7th International Conference on Electrical Information and Communication Technology Eict 2025
dc.relation.urihttps://ieeexplore.ieee.org/document/11355641
dc.subjectLightweight model
dc.subjectSemantic similarity
dc.subjectSkin disease dataset
dc.subjectVision Transformer (ViT)
dc.subjectVision-language models
dc.subjectVisual Question Answering (VQA)
dc.subject.lcshArtificial intelligence--Medical applications.
dc.titleLightweight Visual Question Answering (VQA) model for skin disease detection
dc.typeConference Proceeding
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.identifier.scopus-author-id60522476900
person.identifier.scopus-author-id60522477000
person.identifier.scopus-author-id59011478400
person.identifier.scopus-author-id35108854200
person.identifier.scopus-author-id60522671800
person.identifier.scopus-author-id58660179300
person.identifier.scopus-author-id57222382795

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Demo.pdf.jpg
Size:
2.36 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: