Comparative analysis of attention-based, convolutional, and SSM-based models for multi-domain image classification

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorChakrabarty, Dr. Amitabha
dc.contributor.authorBiswas, Mondrita
dc.contributor.authorRahman, Sayeedur
dc.contributor.authorTarannum, Syeda Farhat
dc.contributor.authorNishanto, Dipro
dc.contributor.authorSafwaan, Md Aqeed
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2025-06-16T10:12:09Z
dc.date.available2025-06-16T10:12:09Z
dc.date.copyright2025
dc.date.issued2025-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 110-116).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2025.en_US
dc.description.abstractThe increasing frequency and severity of environmental and societal challenges, such as natural disasters, medical diagnostics, and agricultural threats require the development of efficient and scalable detection and classification systems. Lightweight and fast models deployed on edge devices, such as surveillance drones, portable diagnostic tools, or agricultural sensors, can address constraints of network delays, adverse conditions, and bandwidth limitations often faced by autonomous technologies. Transformer-based models using attention mechanisms, trade off computational costs to achieve high accuracies in these classification tasks. Recently, State-space models (SSMs) have emerged as a promising alternative in areas where long-range dependence on data is crucial but computational efficiency is particularly important. This research explores the application of Attention-Based, Convolutional, and SSMs, particularly Vision Mamba (ViM), in diverse domains: wildfire detection, plant disease identification, and skin cancer diagnostics. Finally, the feasibility of knowledge distillation in ViM is examined using the information gathered from a thorough evaluation and model comparisons. Evaluations highlight that while CNN models consistently achieved the highest accuracy, ViM Tiny is the most memory-efficient, requiring only 0.03GB of GPU memory. ViM Tiny (7.60M params) achieved 70.60% accuracy in wildfire detection, matching DeiT Base’s (85.80M Params) 70.62% accuracy. The SSM-based models also had the fastest convergence rate. These models achieved promising accuracies in plant disease classification (98.71%–99.65%) and skin cancer detection (87.03%–90.16%), highlighting their potential for efficient and scalable vision tasks. In the context of wildfire detection, knowledge distillation with EfficientNet B7 as a teacher model further improved ViM Tiny’s accuracy from 70.6% to 85.32%, highlighting its potential for lightweight, high-performance applications in critical scenarios.en_US
dc.description.degreeBachelor of Science in Computer Science
dc.description.statementofresponsibilityMondrita Biswas
dc.description.statementofresponsibilitySayeedur Rahman
dc.description.statementofresponsibilitySyeda Farhat Tarannum
dc.description.statementofresponsibilityDipro Nishanto
dc.description.statementofresponsibilityMd Aqeed Safwaan
dc.format.extent117 pages
dc.identifier.otherID: 21101056
dc.identifier.otherID: 21101281
dc.identifier.otherID: 21101016
dc.identifier.otherID: 21101032
dc.identifier.otherID: 21101066
dc.identifier.urihttp://hdl.handle.net/10361/26059
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectState space modelsen_US
dc.subjectVision Mambaen_US
dc.subjectKnowledge distillationen_US
dc.subjectMulti-domain applicationsen_US
dc.subjectAttention-based modelsen_US
dc.subjectConvolutional modelsen_US
dc.subject.lcshNeural networks (Computer science)
dc.titleComparative analysis of attention-based, convolutional, and SSM-based models for multi-domain image classificationen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
21101056, 21101281, 21101016, 21101032, 21101066_CSE.pdf
Size:
9.85 MB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: