Multimodal depression detection from facial emotion dynamics and acoustic features using ensemble learning
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Anwar, Md. Tawhid | |
| dc.contributor.author | Sohan, Md Farhan Mahtab | |
| dc.contributor.author | Shawon, Abu Ferdous | |
| dc.contributor.author | Rahman, Fariha | |
| dc.contributor.author | Morshed, Maimuna | |
| dc.contributor.author | Tithi, Aysha Siddiqua | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-09-14T06:46:54Z | |
| dc.date.available | 2026-09-14T06:46:54Z | |
| dc.date.copyright | 2026 | |
| dc.date.issued | 2026 | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026. | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 98-103). | |
| dc.description.abstract | Depression is a psychiatric disorder that is significantly underdiagnosed in the world, particularly in low-resource countries. This thesis introduces a two-step multimodal framework of automated depression recognition using facial emotion detection. The first phase consists of the systematical training of nine deep learning models CoAtNet-2, ViT-Base/16, Swin-Small, ConvNeXt-Base, MaxViT-T, EfficientNet-B7, ConvNeXt-Small, ResNet- 50, and VGG16 on a carefully cleaned and preprocessed OAHEGA dataset; the top five of them are jointly trained into a weighted soft-voting ensemble, the FER model, with a test accuracy of 93.13. The second step uses the frozen FER ensemble on the D-Vlog depression data, producing per-frame emotion probability sequences that are concatenated with the previously extracted acoustic features on 667 videos that can be used. Using these modalities, a 755 dimensional multi-modal feature representation (250 visual, 500 acoustic, 5 cross-modal features) is created and narrowed down to 195 features using a two-step Mutual Information and RFECV selection pipeline. A number of classifier architectures are compared, such as single base learners, recurrent sequence models and heterogeneous ensembles; a stacking ensemble (MLP, CatBoost, and Random Forest with a Logistic Regression meta-learner) is chosen as the final model, with a test accuracy of 71.81%, and precision of 0.7529 on the held-out fold. The main contributions are a systematic multi-architecture FER comparison, a hybrid ensemble with state-of-the-art accuracy on OAHEGA, a multimodal depression detection pipeline which is the first to use a purposebuilt FER ensemble as a visual backbone, along with audio features, and which can be run on standard consumer hardware to enable scalable, non-invasive screening. | |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Md Farhan Mahtab Sohan | |
| dc.description.statementofresponsibility | Abu Ferdous Shawon | |
| dc.description.statementofresponsibility | Fariha Rahman | |
| dc.description.statementofresponsibility | Maimuna Morshed | |
| dc.description.statementofresponsibility | Aysha Siddiqua Tithi | |
| dc.format.extent | 103 pages | |
| dc.identifier.other | ID 22301272 | |
| dc.identifier.other | ID 22301320 | |
| dc.identifier.other | ID 22301307 | |
| dc.identifier.other | ID 22301319 | |
| dc.identifier.other | ID 22301667 | |
| dc.identifier.uri | https://hdl.handle.net/10361/29906 | |
| dc.language.iso | en_US | |
| dc.publisher | BRAC University | |
| dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 International | en |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/ | |
| dc.subject | Depression | |
| dc.subject | Facial emotion recognition | |
| dc.subject | Hybrid ensemble | |
| dc.subject | Multimodal fusion | |
| dc.subject | D-Vlog | |
| dc.subject | Deep learning | |
| dc.subject.lcsh | Facial expression. | |
| dc.subject.lcsh | Emotion recognition. | |
| dc.subject.lcsh | Depression, Mental. | |
| dc.subject.lcsh | Deep learning (Machine learning). | |
| dc.title | Multimodal depression detection from facial emotion dynamics and acoustic features using ensemble learning | |
| dc.type | Thesis |