Leveraging state-space models for temporal analysis in deepfake detection
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Chakrabarty, Amitabha | |
| dc.contributor.author | Rahman, Affshafee | |
| dc.contributor.author | Bhuiyan, Diniya Tahrin | |
| dc.contributor.author | Khan, Shami Islam | |
| dc.contributor.author | Miah, Salim | |
| dc.contributor.author | Ahsan, MD Shohbat | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-01-12T05:30:11Z | |
| dc.date.available | 2026-01-12T05:30:11Z | |
| dc.date.copyright | 2025 | |
| dc.date.issued | 2025-10 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 60-63). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025. | en_US |
| dc.description.abstract | Deepfake technologies, especially those based on lip-sync forgeries, present an advanced threat to integrity in digital media as they produce seamless audiovisual forgeries that are hard to detect. Transformer-based models show promise, but are resource-heavy and fail to generalize against forgeries created by modern, generative methods. This thesis addresses these issues by proposing an efficient novel framework for the detection of lip-sync forgeries that is based on State-Space Models (SSMs). We propose a dual-stream architecture using parallel Mamba blocks to independently model in the temporal domain the visual dynamics associated with lip movements and the audio dynamics based on audio spectrograms. Both streams use a lightweight MobileNetV3-Small backbone for spatial feature extraction and are configured with an optimal state dimension of 160, discovered through a two-stage ablation study. The resulting temporal feature vectors are fused and a classification is performed using a small MLP head. Trained on the high-quality AV Lips dataset, the Mamba based model proposed achieves a new state of the art accuracy of 94.60% and an AUC of 99.12%, while having an exceptionally low number of parameters, at 2.48 million. In addition, the model achieves robust generalization, emphasizing its potential as a powerful and deployable solution for audio-visual deepfake detection. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Affshafee Rahman | |
| dc.description.statementofresponsibility | Diniya Tahrin Bhuiyan | |
| dc.description.statementofresponsibility | Shami Islam Khan | |
| dc.description.statementofresponsibility | Salim Miah | |
| dc.description.statementofresponsibility | MD Shohbat Ahsan | |
| dc.format.extent | 74 pages | |
| dc.identifier.other | ID 22341024 | |
| dc.identifier.other | ID 22341046 | |
| dc.identifier.other | ID 22301186 | |
| dc.identifier.other | ID 23241051 | |
| dc.identifier.other | ID 22341029 | |
| dc.identifier.uri | http://hdl.handle.net/10361/27422 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Deepfake technologies | en_US |
| dc.subject | Deepfake detection | en_US |
| dc.subject | Lip-sync forgery | en_US |
| dc.subject | State-space models | en_US |
| dc.subject | Multimodal deep learning | en_US |
| dc.subject | Temporal modeling | en_US |
| dc.subject | Audiovisual forgeries | en_US |
| dc.subject | Audiovisual synchronization | en_US |
| dc.subject.lcsh | Deepfakes--Identification. | |
| dc.subject.lcsh | Online manipulation--Prevention. | |
| dc.subject.lcsh | Lipsynching. | |
| dc.title | Leveraging state-space models for temporal analysis in deepfake detection | en_US |
| dc.type | Thesis | en_US |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- 22341024, 22341046, 22301186, 23241051, 22341029_CSE.pdf
- Size:
- 643.39 KB
- Format:
- Adobe Portable Document Format
- Description:
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: