Multimodal depression detection from facial emotion dynamics and acoustic features using ensemble learning

Citation

Abstract

Depression is a psychiatric disorder that is significantly underdiagnosed in the world, particularly in low-resource countries. This thesis introduces a two-step multimodal framework of automated depression recognition using facial emotion detection. The first phase consists of the systematical training of nine deep learning models CoAtNet-2, ViT-Base/16, Swin-Small, ConvNeXt-Base, MaxViT-T, EfficientNet-B7, ConvNeXt-Small, ResNet- 50, and VGG16 on a carefully cleaned and preprocessed OAHEGA dataset; the top five of them are jointly trained into a weighted soft-voting ensemble, the FER model, with a test accuracy of 93.13. The second step uses the frozen FER ensemble on the D-Vlog depression data, producing per-frame emotion probability sequences that are concatenated with the previously extracted acoustic features on 667 videos that can be used. Using these modalities, a 755 dimensional multi-modal feature representation (250 visual, 500 acoustic, 5 cross-modal features) is created and narrowed down to 195 features using a two-step Mutual Information and RFECV selection pipeline. A number of classifier architectures are compared, such as single base learners, recurrent sequence models and heterogeneous ensembles; a stacking ensemble (MLP, CatBoost, and Random Forest with a Logistic Regression meta-learner) is chosen as the final model, with a test accuracy of 71.81%, and precision of 0.7529 on the held-out fold. The main contributions are a systematic multi-architecture FER comparison, a hybrid ensemble with state-of-the-art accuracy on OAHEGA, a multimodal depression detection pipeline which is the first to use a purposebuilt FER ensemble as a visual backbone, along with audio features, and which can be run on standard consumer hardware to enable scalable, non-invasive screening.

Description

This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 98-103).

Publisher Link

Type

Thesis

Creative Commons license

Attribution-NonCommercial-NoDerivatives 4.0 International

Except where otherwise noted, this item's license is described as

Attribution-NonCommercial-NoDerivatives 4.0 International