Multimodal depression detection from facial emotion dynamics and acoustic features using ensemble learning

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorAnwar, Md. Tawhid
dc.contributor.authorSohan, Md Farhan Mahtab
dc.contributor.authorShawon, Abu Ferdous
dc.contributor.authorRahman, Fariha
dc.contributor.authorMorshed, Maimuna
dc.contributor.authorTithi, Aysha Siddiqua
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-14T06:46:54Z
dc.date.available2026-09-14T06:46:54Z
dc.date.copyright2026
dc.date.issued2026
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 98-103).
dc.description.abstractDepression is a psychiatric disorder that is significantly underdiagnosed in the world, particularly in low-resource countries. This thesis introduces a two-step multimodal framework of automated depression recognition using facial emotion detection. The first phase consists of the systematical training of nine deep learning models CoAtNet-2, ViT-Base/16, Swin-Small, ConvNeXt-Base, MaxViT-T, EfficientNet-B7, ConvNeXt-Small, ResNet- 50, and VGG16 on a carefully cleaned and preprocessed OAHEGA dataset; the top five of them are jointly trained into a weighted soft-voting ensemble, the FER model, with a test accuracy of 93.13. The second step uses the frozen FER ensemble on the D-Vlog depression data, producing per-frame emotion probability sequences that are concatenated with the previously extracted acoustic features on 667 videos that can be used. Using these modalities, a 755 dimensional multi-modal feature representation (250 visual, 500 acoustic, 5 cross-modal features) is created and narrowed down to 195 features using a two-step Mutual Information and RFECV selection pipeline. A number of classifier architectures are compared, such as single base learners, recurrent sequence models and heterogeneous ensembles; a stacking ensemble (MLP, CatBoost, and Random Forest with a Logistic Regression meta-learner) is chosen as the final model, with a test accuracy of 71.81%, and precision of 0.7529 on the held-out fold. The main contributions are a systematic multi-architecture FER comparison, a hybrid ensemble with state-of-the-art accuracy on OAHEGA, a multimodal depression detection pipeline which is the first to use a purposebuilt FER ensemble as a visual backbone, along with audio features, and which can be run on standard consumer hardware to enable scalable, non-invasive screening.
dc.description.degreeBachelor of Science in Computer Science and Engineering
dc.description.statementofresponsibilityMd Farhan Mahtab Sohan
dc.description.statementofresponsibilityAbu Ferdous Shawon
dc.description.statementofresponsibilityFariha Rahman
dc.description.statementofresponsibilityMaimuna Morshed
dc.description.statementofresponsibilityAysha Siddiqua Tithi
dc.format.extent103 pages
dc.identifier.otherID 22301272
dc.identifier.otherID 22301320
dc.identifier.otherID 22301307
dc.identifier.otherID 22301319
dc.identifier.otherID 22301667
dc.identifier.urihttps://hdl.handle.net/10361/29906
dc.language.isoen_US
dc.publisherBRAC University
dc.rightsAttribution-NonCommercial-NoDerivatives 4.0 Internationalen
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/4.0/
dc.subjectDepression
dc.subjectFacial emotion recognition
dc.subjectHybrid ensemble
dc.subjectMultimodal fusion
dc.subjectD-Vlog
dc.subjectDeep learning
dc.subject.lcshFacial expression.
dc.subject.lcshEmotion recognition.
dc.subject.lcshDepression, Mental.
dc.subject.lcshDeep learning (Machine learning).
dc.titleMultimodal depression detection from facial emotion dynamics and acoustic features using ensemble learning
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
22301272, 22301320, 22301307, 22301319, 22301667_CSE.pdf
Size:
3.37 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: