Vision transformer-based underwater sound classification of marine mammals
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Chakrabarty, Amitabha | |
| dc.contributor.author | Hasan, Mehedi | |
| dc.contributor.author | Habibullah, Ahmad | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-08-16T05:22:30Z | |
| dc.date.available | 2026-08-16T05:22:30Z | |
| dc.date.copyright | 2025 | |
| dc.date.issued | 2025-12 | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2025. | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 51-52). | |
| dc.description.abstract | Marine mammals rely critically on acoustic communication for survival, yet underwater sound classification faces substantial challenges from complex propagation dynamics, environmental noise, and severe class imbalance in available datasets. This research addresses these challenges by developing and systematically evaluating a comprehensive deep learning framework for marine mammal vocalization classification using the Watkins Marine Mammal Sound Database (WMMSD), comprising 15,567 audio samples across 55 species. We implemented a teacher-student knowledge distillation pipeline, comparing three state-of-the-art architectures—Vision Transformer (ViT), Wave2Vec 2.0, and General Transformer—with the ViT demonstrating superior baseline performance at 89.00% training and 86.85% validation accuracy. Through rigorous data preprocessing including duration filtering, strategic class balancing via SMOTE (Synthetic Minority Oversampling Technique), Mel spectrogram feature extraction (128×224 dimensions), and hybrid augmentation strategies combining SpecAugment and audio-domain transformations, the large ViT teacher model achieved 98.00% training accuracy and 93.85% validation accuracy on both 32-class Best Cut and 55-class All cut subset. A lightweight MobileViTXXS student model (2.1M parameters) was trained via knowledge distillation, retaining 96% of teacher performance at 14× parameter reduction with 84.03% test accuracy and macro-F1 of 0.8402. Post-training optimizations including static quantization and 30% L1-unstructured pruning compressed the model to 1.95 MB with 78.5% accuracy and < 40ms inference latency on edge hardware, enabling real-time deployment on resource-constrained platforms such as Raspberry Pi 4. A practical web application interface was developed for accessible species identification by field researchers. This research contributes methodological innovations in handling severely imbalanced bioacoustics datasets, demonstrates the superiority of knowledge distillation over direct model compression for transformer architectures (avoiding the catastrophic 75% accuracy degradation observed with direct teacher pruning), and establishes a deployment-ready framework for autonomous marine mammal monitoring systems supporting conservation efforts, ship strike prevention, and biodiversity assessment in remote ocean environments. | |
| dc.description.degree | Bachelor of Science in Computer Science | |
| dc.description.statementofresponsibility | Mehedi Hasan | |
| dc.description.statementofresponsibility | Ahmad Habibullah | |
| dc.format.extent | 62 pages | |
| dc.identifier.other | ID 21101062 | |
| dc.identifier.other | ID 21301236 | |
| dc.identifier.uri | https://hdl.handle.net/10361/29129 | |
| dc.language.iso | en_US | |
| dc.publisher | BRAC University | |
| dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 International | en |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/ | |
| dc.subject | Marine bioacoustics | |
| dc.subject | Marine mammals | |
| dc.subject | Sound classification | |
| dc.subject | Knowledge distillation | |
| dc.subject | Deep learning | |
| dc.subject | Class imbalance | |
| dc.subject | Animal communication | |
| dc.subject | Vision transformer | |
| dc.subject | Underwater sound | |
| dc.subject.lcsh | Underwater acoustics. | |
| dc.subject.lcsh | Marine mammals--Vocalization. | |
| dc.subject.lcsh | Sound production by animals. | |
| dc.subject.lcsh | Bioacoustics. | |
| dc.subject.lcsh | Deep learning (Machine learning). | |
| dc.title | Vision transformer-based underwater sound classification of marine mammals | |
| dc.type | Thesis |