Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Leveraging state-space models for temporal analysis in deepfake detection

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorChakrabarty, Amitabha
dc.contributor.authorRahman, Affshafee
dc.contributor.authorBhuiyan, Diniya Tahrin
dc.contributor.authorKhan, Shami Islam
dc.contributor.authorMiah, Salim
dc.contributor.authorAhsan, MD Shohbat
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-01-12T05:30:11Z
dc.date.available2026-01-12T05:30:11Z
dc.date.copyright2025
dc.date.issued2025-10
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 60-63).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.en_US
dc.description.abstractDeepfake technologies, especially those based on lip-sync forgeries, present an advanced threat to integrity in digital media as they produce seamless audiovisual forgeries that are hard to detect. Transformer-based models show promise, but are resource-heavy and fail to generalize against forgeries created by modern, generative methods. This thesis addresses these issues by proposing an efficient novel framework for the detection of lip-sync forgeries that is based on State-Space Models (SSMs). We propose a dual-stream architecture using parallel Mamba blocks to independently model in the temporal domain the visual dynamics associated with lip movements and the audio dynamics based on audio spectrograms. Both streams use a lightweight MobileNetV3-Small backbone for spatial feature extraction and are configured with an optimal state dimension of 160, discovered through a two-stage ablation study. The resulting temporal feature vectors are fused and a classification is performed using a small MLP head. Trained on the high-quality AV Lips dataset, the Mamba based model proposed achieves a new state of the art accuracy of 94.60% and an AUC of 99.12%, while having an exceptionally low number of parameters, at 2.48 million. In addition, the model achieves robust generalization, emphasizing its potential as a powerful and deployable solution for audio-visual deepfake detection.en_US
dc.description.degreeBachelor of Science in Computer Science and Engineering
dc.description.statementofresponsibilityAffshafee Rahman
dc.description.statementofresponsibilityDiniya Tahrin Bhuiyan
dc.description.statementofresponsibilityShami Islam Khan
dc.description.statementofresponsibilitySalim Miah
dc.description.statementofresponsibilityMD Shohbat Ahsan
dc.format.extent74 pages
dc.identifier.otherID 22341024
dc.identifier.otherID 22341046
dc.identifier.otherID 22301186
dc.identifier.otherID 23241051
dc.identifier.otherID 22341029
dc.identifier.urihttp://hdl.handle.net/10361/27422
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectDeepfake technologiesen_US
dc.subjectDeepfake detectionen_US
dc.subjectLip-sync forgeryen_US
dc.subjectState-space modelsen_US
dc.subjectMultimodal deep learningen_US
dc.subjectTemporal modelingen_US
dc.subjectAudiovisual forgeriesen_US
dc.subjectAudiovisual synchronizationen_US
dc.subject.lcshDeepfakes--Identification.
dc.subject.lcshOnline manipulation--Prevention.
dc.subject.lcshLipsynching.
dc.titleLeveraging state-space models for temporal analysis in deepfake detectionen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
22341024, 22341046, 22301186, 23241051, 22341029_CSE.pdf
Size:
643.39 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: