Anwar, Md. TawhidSohan, Md Farhan MahtabShawon, Abu FerdousRahman, FarihaMorshed, MaimunaTithi, Aysha Siddiqua2026-09-142026-09-1420262026ID 22301272ID 22301320ID 22301307ID 22301319ID 22301667https://hdl.handle.net/10361/29906This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.Cataloged from PDF version of thesis.Includes bibliographical references (pages 98-103).Depression is a psychiatric disorder that is significantly underdiagnosed in the world, particularly in low-resource countries. This thesis introduces a two-step multimodal framework of automated depression recognition using facial emotion detection. The first phase consists of the systematical training of nine deep learning models CoAtNet-2, ViT-Base/16, Swin-Small, ConvNeXt-Base, MaxViT-T, EfficientNet-B7, ConvNeXt-Small, ResNet- 50, and VGG16 on a carefully cleaned and preprocessed OAHEGA datasetÍž the top five of them are jointly trained into a weighted soft-voting ensemble, the FER model, with a test accuracy of 93.13. The second step uses the frozen FER ensemble on the D-Vlog depression data, producing per-frame emotion probability sequences that are concatenated with the previously extracted acoustic features on 667 videos that can be used. Using these modalities, a 755 dimensional multi-modal feature representation (250 visual, 500 acoustic, 5 cross-modal features) is created and narrowed down to 195 features using a two-step Mutual Information and RFECV selection pipeline. A number of classifier architectures are compared, such as single base learners, recurrent sequence models and heterogeneous ensemblesÍž a stacking ensemble (MLP, CatBoost, and Random Forest with a Logistic Regression meta-learner) is chosen as the final model, with a test accuracy of 71.81%, and precision of 0.7529 on the held-out fold. The main contributions are a systematic multi-architecture FER comparison, a hybrid ensemble with state-of-the-art accuracy on OAHEGA, a multimodal depression detection pipeline which is the first to use a purposebuilt FER ensemble as a visual backbone, along with audio features, and which can be run on standard consumer hardware to enable scalable, non-invasive screening.103 pagesen-USAttribution-NonCommercial-NoDerivatives 4.0 InternationalBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.http://creativecommons.org/licenses/by-nc-nd/4.0/DepressionFacial emotion recognitionHybrid ensembleMultimodal fusionD-VlogDeep learningFacial expression.Emotion recognition.Depression, Mental.Deep learning (Machine learning).Multimodal depression detection from facial emotion dynamics and acoustic features using ensemble learningThesis