Lightweight visual question answering (VQA) model for skin disease detection
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Chakrabarty, Amitabha | |
| dc.contributor.author | Mahe, Raiyan Habib | |
| dc.contributor.author | Noor, Abtahi | |
| dc.contributor.author | Tasin, Ridwan Noor | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2025-09-02T04:53:30Z | |
| dc.date.available | 2025-09-02T04:53:30Z | |
| dc.date.copyright | 2025 | |
| dc.date.issued | 2025-06 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 73-76). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025. | en_US |
| dc.description.abstract | Visual Question Answering (VQA) is an area of artificial intelligence that combines image analysis with natural language understanding to generate context-aware answers to visual queries. Its application in the medical field, particularly in dermatology, holds transformative potential by enabling accessible, interpretable, and efficient diagnostic support. However, there is a lack of structured VQA dataset for skin diseases that can be used for training as well as benchmarking models. Moreover, existing VQA model are all extremely heavy weight and require specialized hardware to train and run. Hence, we developed a custom dataset of 1,038 images annotated by a medical expert for 11 disease classes, with seven question-answer pairs per image. The entire dataset is split into three sections: Train set with 833 images, Test set with 103 images and Validation set with 102 images. In addition, this research proposes a lightweight VQA model pipeline capable of identifying common skin diseases from images and responding to clinically relevant questions related to disease name, severity, causes, diagnostic approach, prevention, contagiousness, and cancer risk. The model uses a modular architecture that integrates a Vision Transformer (ViT) with 86 million parameters for image encoding and MiniLM, a transformer-based text encoder with 22 million parameter, ensuring high accuracy while minimizing computational requirements. We have also trained state-of-theart vision language models such as Gemma-3, QwenVL-2.5, LLaVA-1.5, and BLIP-2 using our dataset for comparison. We achieved high semantic similarity scores on these models having 93.52%, 94.54%, 93.78%, and 94.47% BertScore on Gemma-3, QwenVL-2.5, LLaVA-1.5, and BLIP-2. Unlike these models, our solution is optimized for performance in low-resource settings. The system provides an intuitive interface for users to interact and receive actionable insights, benefiting both patients and healthcare professionals by reducing diagnostic delays and improving decisionmaking with accuracy of 94.87% with a total parameter count of only 108 million, rivaling those with billions of parameters. The results demonstrate that lightweight, domain-adapted VQA models can effectively bridge the healthcare accessibility gap through accurate and interpretable AI assistance. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Raiyan Habib Mahe | |
| dc.description.statementofresponsibility | Abtahi Noor | |
| dc.description.statementofresponsibility | Ridwan Noor Tasin | |
| dc.format.extent | 76 pages | |
| dc.identifier.other | ID 21301737 | |
| dc.identifier.other | ID 21301304 | |
| dc.identifier.other | ID 21101326 | |
| dc.identifier.uri | http://hdl.handle.net/10361/26630 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Visual question answering | en_US |
| dc.subject | Medical imaging | en_US |
| dc.subject | Artificial intelligence | en_US |
| dc.subject | Disease detection | en_US |
| dc.subject | Natural language processing | en_US |
| dc.subject.lcsh | Natural language processing (Computer science). | |
| dc.subject.lcsh | Visual Perception. | |
| dc.subject.lcsh | Skin--Diseases--Treatment. | |
| dc.subject.lcsh | Artificial intelligence. | |
| dc.title | Lightweight visual question answering (VQA) model for skin disease detection | en_US |
| dc.type | Thesis | en_US |