Cross-attention fusion vision transformer for explainable and efficient multi-class eye disease detection
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Rasel, Annajiat Alim | |
| dc.contributor.advisor | Agomoni, Ahmed Mayeesha Reza | |
| dc.contributor.author | Rahman, Yasin | |
| dc.contributor.author | Mozahedul Hoque, Md. | |
| dc.contributor.author | Chowdhury, Moriyum | |
| dc.contributor.author | Mustakim Al Mahmud | |
| dc.contributor.author | Mahi, Mahidul Islam | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-04-22T06:56:13Z | |
| dc.date.available | 2026-04-22T06:56:13Z | |
| dc.date.copyright | 2026 | |
| dc.date.issued | 2026-01 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 91-95). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026. | en_US |
| dc.description.abstract | Early and accurate detection of retinal fundus diseases is critical for preventing irreversible vision loss and supporting effective clinical decision-making. Retinal fundus imaging is widely used for large-scale screening due to its non-invasive nature; however, the diversity and structural complexity of retinal pathologies pose significant challenges for automated analysis. Convolutional Neural Networks (CNNs) have been extensively employed for fundus image classification owing to their strong local feature extraction capabilities, however, their limited ability to model long-range contextual dependencies constraints performance in complex multi-class disease scenarios. Vision Transformers (ViTs), on the other hand, leverage self-attention mechanisms to capture global contextual information but often suffer from high computational costs and reduced effectiveness in limited-data medical imaging settings. The study proposed a lightweight hybrid Cross-Attention Fusion Vision Transformer architecture can be used to classify multi-class retinal diseases through fundus images. The suggested model combines CNN-based local feature extraction and transformerbased global contextual modelling with the cross-attention fusion mechanism, which allows both fine-grained pathological features and holistic retinal structure to interact and at the same time keep computational efficiency. The hybrid model is specifically designed with small-scale medical data in mind and uses attention-based interpretability, which helps to include explainable AI in the future to improve clinical transparency. Experimental evaluation on publicly available fundus datasets demonstrates that the proposed hybrid approach achieves a favorable balance between accuracy, robustness, and efficiency compared to standalone CNN and Vision Transformer models, highlighting its suitability for automated retinal disease screening applications. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Yasin Rahman | |
| dc.description.statementofresponsibility | Md. Mozahedul Hoque | |
| dc.description.statementofresponsibility | Moriyum Chowdhury | |
| dc.description.statementofresponsibility | Mustakim Al Mahmud | |
| dc.description.statementofresponsibility | Mahidul Islam Mahi | |
| dc.format.extent | 95 pages | |
| dc.identifier.other | ID 23341115 | |
| dc.identifier.other | ID 22101379 | |
| dc.identifier.other | ID 21301210 | |
| dc.identifier.other | ID 24241323 | |
| dc.identifier.other | ID 21301542 | |
| dc.identifier.uri | http://hdl.handle.net/10361/28029 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Retinal fundus imaging | en_US |
| dc.subject | Eye disease classification | en_US |
| dc.subject | Convolutional neural networks | en_US |
| dc.subject | Vision transformer | en_US |
| dc.subject | Lightweight deep learning | en_US |
| dc.subject.lcsh | Ophthalmology--Data processing. | |
| dc.subject.lcsh | Deep learning (Machine learning). | |
| dc.subject.lcsh | Eye--Diseases--Diagnosis--Data processing. | |
| dc.subject.lcsh | Diagnostic imaging--Computer-aided design. | |
| dc.subject.lcsh | Diagnostic imaging--Digital techniques. | |
| dc.title | Cross-attention fusion vision transformer for explainable and efficient multi-class eye disease detection | en_US |
| dc.type | Thesis | en_US |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- 23341115, 22101379, 21301210, 24241323, 21301542_CSE.pdf
- Size:
- 1.38 MB
- Format:
- Adobe Portable Document Format
- Description:
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: