Thesis (Bachelor of Science in Computer Science and Engineering)

Permanent URI for this collectionhttps://hdl.handle.net/10361/28433

Browse

Recent Submissions

Now showing 1 - 20 of 1043
  • listelement.badge.dso-type Item ,
    Open Access
    An interpretable KAN-guided SNN architecture with channel-wise contribution tracking for EEG-based emotion recognition
    (BRAC University, 2026-06) Kamal, Abdul Kariem Bin; Fardin, S.M. Hasnat; Maisha, Sumaiya Khan; Samin, Saeeb Rahman; Alam, Md. Ashraful; Alam, Md. Golam Rabiul; Department of Computer Science and Engineering
    EEG-based emotion recognition is challenging due to high signal-to-noise ratio, subtle difference in emotional states and mainly inter-participant variability. Previous works have shown that Spiking Neural Networks (SNN) are suitable for capturing temporal EEG activity, however, their learned representations are difficult to interpret at a channel level. This study proposes a channel-wise KAN (Kolmogorov- Arnold Networks) guided SNN for interpretable emotion recognition using EEG data from the DER-VREEG dataset. The dataset classifies emotion from four distinct emotion classes: happy, bored, calm and scared using four channels that capture EEG data from TP9, TP10, AF7 and AF8 which are four distinct regions of the human brain. Unlike traditional SNN classifiers, our proposed model can estimate the contribution factor of each of the four EEG channels towards predicting a certain emotion class, which is crucial for neuroscientists to study how a specific region of the brain reacts to a specific emotion. The framework thus provides explainable EEG emotion recognition along with classification accuracy that is on par with existing related architectures.
  • listelement.badge.dso-type Item ,
    Open Access
    Multimodal depression detection from facial emotion dynamics and acoustic features using ensemble learning
    (BRAC University, 2026) Sohan, Md Farhan Mahtab; Shawon, Abu Ferdous; Rahman, Fariha; Morshed, Maimuna; Tithi, Aysha Siddiqua; Anwar, Md. Tawhid; Department of Computer Science and Engineering
    Depression is a psychiatric disorder that is significantly underdiagnosed in the world, particularly in low-resource countries. This thesis introduces a two-step multimodal framework of automated depression recognition using facial emotion detection. The first phase consists of the systematical training of nine deep learning models CoAtNet-2, ViT-Base/16, Swin-Small, ConvNeXt-Base, MaxViT-T, EfficientNet-B7, ConvNeXt-Small, ResNet- 50, and VGG16 on a carefully cleaned and preprocessed OAHEGA dataset; the top five of them are jointly trained into a weighted soft-voting ensemble, the FER model, with a test accuracy of 93.13. The second step uses the frozen FER ensemble on the D-Vlog depression data, producing per-frame emotion probability sequences that are concatenated with the previously extracted acoustic features on 667 videos that can be used. Using these modalities, a 755 dimensional multi-modal feature representation (250 visual, 500 acoustic, 5 cross-modal features) is created and narrowed down to 195 features using a two-step Mutual Information and RFECV selection pipeline. A number of classifier architectures are compared, such as single base learners, recurrent sequence models and heterogeneous ensembles; a stacking ensemble (MLP, CatBoost, and Random Forest with a Logistic Regression meta-learner) is chosen as the final model, with a test accuracy of 71.81%, and precision of 0.7529 on the held-out fold. The main contributions are a systematic multi-architecture FER comparison, a hybrid ensemble with state-of-the-art accuracy on OAHEGA, a multimodal depression detection pipeline which is the first to use a purposebuilt FER ensemble as a visual backbone, along with audio features, and which can be run on standard consumer hardware to enable scalable, non-invasive screening.
  • listelement.badge.dso-type Item ,
    Open Access
    Real-time Bangladeshi sign language recognition and natural language interpretation using hybrid deep learning and LLMs
    (BRAC University, 2026-06) Fairuz, Fabliha Akther; Roy, Santonu; Mahmud, Sajid; Tasnim, Sumiya; Rafey, Inkiad Bin Ershad; Alam, Md. Ashraful; Department of Computer Science and Engineering
    Deaf and hard-of-hearing communities in Bangladesh face a persistent communication gap: existing sign language recognition systems process hand gestures but ignore the facial expressions and head movements that carry grammatical meaning in Bangladeshi Sign Language (BdSL). This study introduces BdSL-NMM, the first BdSL dataset annotated with non-manual marker labels alongside word-level glosses, covering 62 sign classes and 5 expression categories across 4,170 recordings from seven native signers. To evaluate on this dataset, a multi-stream Transformer architecture called SignNet-V2 was developed, which simultaneously recognises sign words and non-manual expression markers from skeletal landmark sequences. The system processes four input modalities, namely body pose, left hand, right hand, and face mesh, through dedicated stream-specific encoders, cross-stream attention fusion, hierarchical temporal encoding, and a multi-task classification head. All models were trained and evaluated under leave-one-signer-out cross-validation, where the test signer never appeared during training, providing a realistic measure of cross-signer generalisation. Under this protocol, SignNet-V2 achieves 38.44% top-1 word recognition accuracy. For expression recognition, a 56-dimensional framework of normalised geometric and temporal features achieves 66.39% overall accuracy under signer-independent evaluation, with neutral and negation recall of 97.22% and 91.18% on the unseen test signer. A second interpretation stage passes recognised signs and expression tags to Google Gemini, which generates grammatically correct Bengali sentences. This work establishes the first published benchmarks for simultaneous word and non-manual marker recognition in BdSL under signer-independent evaluation.
  • listelement.badge.dso-type Item ,
    Open Access
    MIM-HeMIS: Self-supervised masked image modeling with heterogeneous modality fusion for brain tumor segmentation under missing MRI modalities
    (BRAC University, 2026-06) Hasan, Md. Jahidul; Opy, Raduan Ahmed; Shafi, Md Alif Khan; Jameel, Ahnaf Ashraf; Alvi, Md. Jobayer; Ahmed, Md. Sabbir; Department of Computer Science and Engineering
    Accurate segmentation of regions of interest from MRI to diagnose brain tumors is crucial. MRI is the standard technique used in imaging the brain, but manual segmentation of tumor regions is time-consuming, manual and prone to human error. Although deep learning has been able to significantly enhance segmentation performance, it requires big data sets of labeled images which is difficult in medical applications. In this study, we introduce a hybrid CNN-Transformer framework that integrates the Masked Image Modeling (MIM) based self-supervised pretraining and the HeMIS statistical abstraction mechanism. A hybrid model is pretrained on the BraTS 2021 dataset and finetuned for 4-class brain tumor segmentation for all 4 MRI modalities following the proposed framework. A major advantage is that the modalities are robust, meaning that there is no need to train a new model for each of the 15 possible combinations of modalities. When modalities are missing, the variance computed by the HeMIS abstraction layer captures feature disagreement across available inputs, providing useful information for identifying cases where pre-dictions may be less reliable. Though this variance has not been calibrated as a clinical confidence score. The model was tested on BraTS 2021 with all 15 modality combinations with excellent and clinically applicable results: Whole Tumor Dice of 90.2%, Tumor Core Dice of 85.9%, and Enhancing Tumor Dice of 79.1% for all four modalities.
  • listelement.badge.dso-type Item ,
    Open Access
    JASAL: A remote patient health monitoring system driven by federated learning
    (BRAC University, 2026) Alif, Md. Labibul Ahsan; Ahmed, Juboraj; Ahmed, Shajid; Lubaba, Nafisa; Islam, Ayasha; Mostakim, Moin; Department of Computer Science and Engineering
    Remote patient monitoring (RPM) is growing fast, but most systems today rely on centralizing sensitive health data in the cloud. This raises serious concerns: patient privacy, strict regulations (such as GDPR and HIPAA), and “data silos,” in which hospitals can’t learn from one another without sharing private records. To address this, this thesis introduces JASAL—a privacy-preserving federated learning framework for real-time anomaly detection in distributed healthcare networks. Instead of moving patient data to a central server, JASAL trains small, lightweight Deep Autoencoders directly on edge devices (like IoT gateways at hospitals or clinics). Only encrypted model updates, not raw patient data, are sent to a coordinator for aggregation. The Hybrid Differential Privacy – Secure Multi-party Computation (DP-SMPC) of the JASAL system provides a dual level of privacy protection. The differential privacy level (ϵ = 1.0) provides protection from the published output of the learning task by concealing individual contributions of all patients participating; and the secure multi-party computation (SMPC) provides confidentiality of the models’ updates by protecting the contributions of each hospital from all other hospitals in the group. This defense-in-depth approach to physician and patient privacy ensures sensitive patient information will remain safe without sacrificing operational performance. In testing, JASAL performed well on a realistic synthetic dataset of 32,796 patients from 10 heterogeneous hospital nodes, achieving 96.28% overall diagnostic accuracy and an AUC-ROC of 0.978 with regard to its ability to reliably function in the detection of complex clinical conditions, such as sepsis, across multiple diverse clinical patient populations. Further, JASAL also exhibited a rather remarkable 98.3 percent reduction in bandwidth costs (from 20 MB to 1.2 GB) from transmitting raw data centrally to the hospitals while providing linear scalability (R² = 0.999) and ultra-low-inference latency (15.6 ms) sufficient for bandwidth-constrained, rural hospital networks. Tackling trust and usability of clinical artificial intelligence, JASAL’s lightweight Vectorized Attribution Engine produces very user-friendly feature level explanations (for example, ”Sepsis Pattern Detected: rising heart rate + falling SpO2”) in under 0.82 milliseconds, helping providers understand the basis of an alert. In addition, JASAL’s Adaptive Personalization Module automatically personalizes patientspecific thresholds to reduce alarm fatigue and to minimize false positives by as much as 49% compared to traditional static systems. Overall, JASAL demonstrates that strong privacy and high performance can coexist in real-world medical systems. It offers a practical, ethical, and scalable path toward truly autonomous, privacy-by-design remote patient monitoring.
  • listelement.badge.dso-type Item ,
    Open Access
    HALP: A hybrid anonymous login protocol
    (BRAC University, 0031-08) Tonay, Zunaed Sazzad; Ajmotgir, Elham M.; Ferdous, Md Sadek; Department of Computer Science and Engineering
    The current research demonstrates the design and implemention of the Hybrid Anonymous Login Protocol (HALP), which attempts to overcome the main shortcomings of existing privacy-preserving authentication systems. Reliance on fixed identifiers in conventional authentication systems implies that the attributes associated with a user’s privacy are exposed to tracking and profiling threats. HALP combines Zero-Knowledge Proofs (ZKPs), pseudonymous identifiers and selective revocation protocols to accomplish unlinkability, scalability, and privacy protection. The ability to log in to the system without revealing identifiable information allows the user to maintain their privacy, while selective revocation provides the ability to prevent compromised credentials from gaining access to services. HALP is designed modularly and conforms to decentralized identity specifications (most notably W3C Verifiable Credentials), enabling its use in privacy-oriented applications such as healthcare, e-voting and decentralized applications (DApps). The framework’s performance and scalability are thoroughly evaluated through benchmarking tests, and security is analyzed in terms of cryptographic properties and security games, through which HALP was found to be effective for deployment in real-life scenarios
  • listelement.badge.dso-type Item ,
    Open Access
    AI phishing detection tool
    (BRAC University, 2026-01) Udatta, Asif Ad-Deen; Nahid, Naimur Rahman; Sifat, Shafayet Noor; Anoy, Raiyan Rafiz; Miraj, Sayeed Bin; Zahid, Imran; Department of Computer Science and Engineering
    Phishing attacks remain one of the most critical cybersecurity threats, exploiting malicious URLs to deceive users and extract sensitive information. Traditional detection approaches, such as blacklist-based and rule-based systems, are increasingly ineffective against rapidly evolving and previously unseen phishing techniques. To address this challenge, this study proposes a hybrid deep learning framework that integrates a supervised transformer-based model (RoBERTa) with an unsupervised Autoencoder for anomaly detection using structured URL features. The RoBERTa model captures contextual and semantic patterns from raw URLs, while the Autoencoder identifies deviations through reconstruction error, enabling detection of unknown phishing behaviors. The framework is evaluated using standard metrics including accuracy, precision, recall, F1-score, and AUC, along with cross-validation to ensure robustness. Experimental results demonstrate that the hybrid model achieves near-perfect performance, significantly outperforming individual approaches while maintaining strong generalization capability. The combination of supervised and unsupervised learning enhances detection reliability and reduces false positives. Overall, this research provides an effective, scalable, and adaptive solution for real-world phishing detection in dynamic cybersecurity environments.
  • listelement.badge.dso-type Item ,
    Open Access
    Understanding food selection behavior using HCI-driven smart menu interfaces
    (BRAC University, 2026-01) Zaman, Rahmin Amer; Kabir, Saiara; Khan, MD Sidratul Muntaha; Rahman, MD Rashadat Abdullah; Dofadar, Dibyo Fabian; Department of Computer Science and Engineering
    Food waste in University cafeterias has been a persistent sustainability challenge which is usually fuelled by the imbalance between the food choice customer expectation and consumption. The current interventions are mostly oriented at the postconsumption monitoring, which does not emphasize the most important moment of food selection. The method adopted in this research is that of Human-Computer Interaction (HCI) with the view to understand how smart menu system can be used to intervene at the point of decision to reduce food wastage. Following the previous prototype studies, this phase of the research uses qualitative and prototype-based assesment using surveys with 50 respondents and semi-structured interviews with 20 stakeholders, including students, faculty members, cafeteria staff, and management. The subjects were manipulated using a set of 3 menu paradigms, a conventional listing-based menu with a relatively simple set of nutritional information, information-lite menu with a simple nutritional information, and information-rich smart menu with visual previews, sustainability related cues and waste related information. Thematic analysis of the interviews shows that food waste is often caused as a result of expectation-reality discrepancies, time-stressed decisions, social factors, and no control over the amount of portion. Participants constantly emphasised the need of clarity in visuals and belief in quality of food and plainness of design in facilitating improved decisions. Altogether, the results demonstrate the potential of smart menus that are HCI-based, low cost and behaviorally informed, to be able to impact the food preferences prior to eating, and still fitting in with the current cafeteria operations. The research contributes qualitative knowledge to the domain of sustainable interaction design and identifies the directions of practical implementation and testing in the future.
  • listelement.badge.dso-type Item ,
    Open Access
    Multi-architecture deep learning framework for glaucoma detection using CNN and vision transformer models
    (BRAC University, 2026-01) Tahrim, Syeda Tauba; Islam, Maiesha; Podder, Ayon; Tasnim, Nishat; Rahman, Rafeed; Department of Computer Science and Engineering
    Glaucoma as a primary cause of permanent blindness is still unnoticeable in the initial stages due to the absence of symptoms. The purpose of this study is to design an automated system of glaucoma detection, which will be based on the use of deep learning algorithms to retinal fundus images. The proposed work has introduced a new image preprocessing methodology that will involve the combination of Contrast Limited Adaptive Histogram Equalization (CLAHE), gamma correction, and green channel extraction in order to enhance the quality of the image and make the bright points including optic nerve head and retinal vasculature more discernible and visible that is vital in the detection of glaucoma correctly. The efficacy of four various deep learning structures, such as ResNet50, EfficientNetB0, SwinTransformer, and Graph Convolutional Networks (GCNs) in glaucoma detection, have been tested. In the present study, the performance of four deep learning models, namely ResNet50, EfficientNetB0, SwinTransformer, and Graph Convolutional Networks (GCNs), is evaluated to determine the effectiveness of the models in the detection of glaucoma. Among the four models, the ResNet50 model recorded the highest performance, with an accuracy of 99.47%, sensitivity of 100%, and specificity of 99.33%. In the present study, the application of the ResNet50 model in the clinical domain is also explored. For the application, the model is converted into the pytorch format, which is compatible with all platforms. It is also optimized to run the model efficiently. A streamlit interface is developed to deploy the ResNet50 model in the clinical domain for the detection of glaucoma. In the Gradio interface, the clinicians can upload the images of the retinal fundus, and the model will give instant results. In the proposed system, the cloud, edge, and hybrid architectures are used. Nevertheless, the nextgeneration study must focus on the multi-center validation, which will assess the generalizability of the offered model. The applicability and credibility of the proposed model will be enhanced by the addition of other diagnostic tools, including Optical Coherence Tomography (OCT), and the development of eXplainable Artificial Intelligence (XAI) models. The suggested study demonstrates the feasibility of deep learning application in the glaucoma detection, and it offers an effective solution to the early diagnosis and treatment of the disease.
  • listelement.badge.dso-type Item ,
    Open Access
    Accessible data visualization for ASRS-positive university students: A multi-criteria analysis of performance efficiency and user engagement
    (BRAC University, 2024) Shovon, Md Rakibur Rahman; Khan, Rafiad Zaman; Islam, Kazi Wahidul; Tamim, Rifat Mahmud; Mukta, Jannatun Noor; Department of Computer Science and Engineering
    Attention-Deficit Hyperactivity Disorder (ADHD) creates substantial educational and economic challenges, but digital information architecture often overlooks neurodiverse cognitive profiles, which results in overcrowded designs that cause cognitive overload. The study focuses on the effectiveness of the various data visualization formats on the processing efficiency and physiological stress among University students classified by the Adult ADHD Self-Report Scale (ASRS v1.1). The methodology employed a multi-criteria experimental study with 70 participants (40 ASRSnegative, 30 ASRS-positive) evaluating eight different visualization formats: Plain Text, Highlighted Text, Bar Charts, Bar Charts with Images, Pictographs, Simplified Infographics, Detailed Infographics, and Text modified by Artificial Intelligence. Inverse Efficiency Score (IES) was used to measure performance, and medical-grade pulse oximeters recorded real-time physiological stress measurements. Statistical rigour was ensured using Two-Way Mixed ANOVA, Linear Mixed-Effects Models (LMM), MANOVA, Bonferroni correction and Bayes Factors. Results revealed a very strong level of interaction (p < 0.001) that visualization formats have different effects on individuals with ADHD or ADHD likelihood. Plain Text, Bar Charts, and Bar Charts with Images maintained almost similar cognitive equity between both groups. Contrarily, Detailed Infographics and AI-Modified Text caused severe performance impairment (Hedges’ g = -1.775 and -1.162, respectively). Paradoxically, traditional attentional aids like text highlighting acted as visual noise, impairing ASRS-positive performance. Moreover, a critical finding revealed preferenceperformance dissociation: ADHD and ADHD-likelihood (ASRS-positive) individuals favoured visually complex designs that objectively hindered their accuracy and processing speed, underscoring the necessity for empirical metrics over subjective feedback. The practical design suggestions presented in the findings are based on concrete, empirically grounded results to establish digital learning environment settings that accommodate ADHD and ADHD-likelihood population.
  • listelement.badge.dso-type Item ,
    Open Access
    TEBP2-YOLOv8: A defect-sensitive convolutional neural network architecture featuring texture enrichment and P2 micro-scale detection for fine-grained eggplant quality assessment in precision agriculture
    (BRAC University, 2026-04) Islam, S. M. Ababil; Saha, Debashis; Alam, Md. Ridowanul; Reza, Md. Tanzim; Department of Computer Science and Engineering
    Bangladesh’s economy is mostly dependent on agriculture. Ensuring quality control of fruits plays a crucial role in the agricultural industries of countries like Bangladesh. Eggplant is one of the most popular vegetables in Bangladesh. Detecting and classifying eggplant fruit surface defects are important for maintaining the quality of the eggplant. However, the eggplant quality assessment still relies on traditional methods, which are less accurate, time-consuming, and leads to human error. To overcome these challenges, we propose a deep learning based approach using Convolutional Neural Networks (CNNs) for the classification of eggplant fruit quality based on the condition of defects. A total of 7,686 images of different varieties of eggplants with defects from local markets and agricultural fields of Dhaka, Bogura and Satkhira districts of Bangladesh have been collected to make the dataset. This dataset has been used to train our CNN models to classify the eggplants into three categories - rotten, insect damage, and physical damage. We have initially used YOLOv8x, YOLOv8s, Faster R-CNN for training the dataset. Later, we have proposed a custom YOLOv8x model named TEBP2-YOLOv8x having an additional P2 head for improved smaller defect detection and Texture Enhanced Blocks (TEB) for texture enhancement of subtle defects. We have also used Weighted Boxes Fusion (WBF) for combining our proposed TEBP2-YOLOv8x and YOLOv8s for improved precision. Various data augmentation techniques are also used for better accuracy. After evaluating the trained models on test set images, the results demonstrate that our proposed TEBP2-YOLOv8x model has the best performance (mAP@0.5=0.924) among the other models. An ablation study on TEBP2-YOLOv8x model shows that it enhances the detection accuracy of insect class with tiny and subtle defects (mAP@0.5 increases from 0.836 to 0.904). We believe that this study can contribute to automate the eggplant defect classification with enhanced accuracy and it will reduce the economic losses in the agricultural sector by ensuring consistent eggplant fruit quality.
  • listelement.badge.dso-type Item ,
    Open Access
    FeastAl: An ML & LLM-powered dinner selection web application
    (BRAC University, 2026-04) Mukul, Rifat Mahamud; Ahmed, Saadat Rafid; Department of Computer Science and Engineering
    In today’s fast-paced Restaurant Industry of Bangladesh, discovering the perfect diner has become an experience that goes beyond the basic nature of searching; that is why a personalized and intelligent recommendation system has become essential to help Bangladeshi users to come up with an intuitive and e!ective way of searching restaurants. This project introduces a comprehensive restaurant recommendation engine that uses advanced machine learning algorithms and large language models (LLMs) to find the best eating options for each user. By assessing crucial characteristics such as geographical proximity, the system narrows down options based on the user’s distance from probable restaurants, assuring convenience. After that, it uses sophisticated sentiment analysis and rate evaluation algorithms to analyze restaurant reviews, providing unbiased information on the caliber of the cuisine and the level of service. Furthermore, the engine incorporates user history and behavior tracking, learning from previous choices and dining patterns to recommend restaurants that match individual tastes. The unique time-based rating feature of this project ensures that suggestions are both enticing and useful by balancing the user’s available eating time with the restaurant’s food preparation time. In order to customize recommendations to each user’s unique time limitations and culinary tastes, the system also looks into connections between food type and preparation time. The ultimate result is a strong, data-driven recommendation system that dynamically aligns high-quality dining experiences with both personal preferences and practical logistical issues.
  • listelement.badge.dso-type Item ,
    Open Access
    PAUSE: Post-hoc sequential instance unlearning in panoptic segmentation via asymmetric conflict aware gradient projection
    (BRAC University, 2026-06) Ahammed, Moin; Adib, Md. Azmain; Islam, Md. Imdadul; Alam, Md. Golam Rabiul; Department of Computer Science and Engineering
    Panoptic segmentation models face increasing pressure to comply with privacy regulations and data-subject withdrawals, necessitating the post-hoc removal of specific training data. While machine unlearning bypasses the prohibitive cost of full retraining, it remains severely underexplored in dense prediction tasks. This challenge is particularly acute in transformer-based decoders, which couple individual object instances through shared query representations; consequently, suppressing a single target risks catastrophic degradation of its entire semantic class. This bottleneck intensifies under sequential deletion, where continuous updates often accumulate collateral damage or restore previously forgotten targets. To address this, we propose PAUSE (Post-Hoc Asymmetric Unlearning for Sequential Erasure), a post-hoc teacher-student framework tailored for query-structured architectures like Mask2Former. By integrating localized instance erasure with asymmetric gradient projection, damage-aware dynamic weighting, and a priority-ordered replay memory, PAUSE safely untangles the unlearning objective from shared features. Extensive evaluations on the Cityscapes dataset against NegGrad, NegGrad+, and SCRUB baselines demonstrate that PAUSE successfully processes sequential deletion queues, cleanly erasing targeted instances while bounding global utility degradation to a marginal ΔPQ of −0.4 percentage points. Finally, we formalize the fundamental tension between class-level retention and instance-level erasure, establishing a critical trade-off baseline for future dense-prediction unlearning systems.
  • listelement.badge.dso-type Item ,
    Open Access
    End-to-end pipeline: Normalization, and summarization of Bangla-English code-switching conversation
    (BRAC University, 2026-01) Rahman, Samir; Siddique, Dania; Tasnim, Humaira Sadia; Khan, Zahidul Islam; Omar, Nayem Bin; Islam, Nazmul; Department of Computer Science and Engineering
    Conventional Natural Language Processing (NLP) systems are predominantly designed and trained for monolingual text. However, the extensive use of Bangla-English and Banglish( Bengali written in Romanized alphabets) code-switching informal conversations in digital communication proves to be challenging for these NLP systems. To addresses this research gap in processing mixed language texts of Bangla-English-Banglish we proposed BiLoRA-BN, a end-to-end pipeline specifically designed for normalization and summarization of code-switched Bangla-English-Banglish conversations in digital communication. The proposed system employs a two-stage Low-Rank Adaptation (LoRA) architecture built on a shared, pre-trained Transformer as backbone with 4-bit quantization, enabling efficient multi-task learning while reducing trainable parameters. The experimental results of BiLoRA-BN are compelling, it significantly outperforms conventional sequential and cascading pipelines, with a +5.74 BLEU gain in normalization quality and a +5.26 ROUGE-1 improvement in final summary accuracy. During interface testing BiLoRA-BN also delivers results faster compared to other pipeline based cascading approach of different architecture and pre-trained models. Crucially, in the interface part, the entire system of BiLoRA-BN can operates on a consumer-grade GPUs with 8GB of memory. By directly modeling all the transformations BiLoRA-BN tries to capture the reality of multilingual digital discourse, with the complex scenario like code-switching and code-mixing in the conversations. This work contributes a step toward understanding how people naturally speak and write to communicate in digital spaces and how NLP model work with it.
  • listelement.badge.dso-type Item ,
    Open Access
    Improving structural code quality of java source codes through automated detection and refactoring of code smells leveraging large language models (LLMs)
    (BRAC University, 2025) Daraksha, Tahura Alam; Sikder, Hironmoy; Esha, Khadija Farhana; Kamran, Fardin; Zawhar, Ganim; Azmain, Md. Aquib; Rifat, Riazul Islam; Department of Computer Science and Engineering
    Structural and design issues in source code, commonly referred to as code smells, can reduce software maintainability, readability, and overall quality. Traditional code analysis tools rely primarily on predefined rules and heuristics for smell detection, which often limits their ability to understand broader code context and provide meaningful refactoring support. Recent advances in Large Language Models (LLMs) have created new opportunities for intelligent code analysis and automated software maintenance. This thesis presents an LLM-based approach for the detection and refactoring of code smells in source code. The proposed framework employs two specialized models: a code smell detection model trained to identify common structural and design issues, and a refactoring model trained on paired before-and-after code examples to generate refactored code and identify the refactoring technique applied. The detection model recognizes several categories of code smells, while the refactoring model generates improved code by applying appropriate refactoring transformations. We are working to propose an automated engine that uses LLM to detect and refactor key code smells in deployed Java projects. Our proposed framework combines static code heuristics with parameter-efficient fine-tuning using for code smell detection. Additionally, a separate LoRA fine-tuned model is trained on before-and-after code pairs to perform refactoring , generating refactored code along with the corresponding refactoring type. We aim to find if the semantic understanding of LLMs are capable enough to find the common smells in deployed Java source codes and refactor to improve the quality of the codes. We will construct a dataset for internal testing and benchmarking. This work illustrates how LLMs can serve as intelligent assistants in software maintenance, enabling automated detection and context-aware refactoring.
  • listelement.badge.dso-type Item ,
    Open Access
    Effects of excessive PUBG gaming hours on educational performance and mental health of university students in Bangladesh: Evidence from SPSS and SmartPLS
    (BRAC University, 2026-04) Tahana, Ummul Mum; Khan, Tammim Liza; Khan, Fahima Hasin; Samadder, Atoshi; Tan, Tamkin Mahmud; Khan, A S M Nasim; Department of Computer Science and Engineering
    PUBG is a famous game for the young generation among online video streaming game. PUBG gaming addiction has raised concern regarding academic performance and mental health. The aim of this research is to study the effects of PUBG gaming intensity on academic performance, mental health, sleep quality, social engagement. A questionnaire survey is conducted to collect data from 523 undergraduate students of both Public and Private universities of Bangladesh. For statistical analysis SPSS software is used and the SEM model is constructed through SmartPLS for analyzing relationships among variables. The findings show that excessive PUBG gaming highly impacts academic performance. The students who spend more time gaming suffer from stress anxiety, poor sleep quality and academic disruption. All these factors moderate the CGPA of the student. Moreover, PUBG players become gradually involved in gaming and interact with other players online which leads to social engagement in gaming circles. More gameplay finally results in less CGPA. This research highlights the interrelation among the constructs including gaming intensity, CGPA, stress and anxiety, sleep quality and social engagement from the software suggesting the importance of balanced gaming habits for academic and mental well being.
  • listelement.badge.dso-type Item ,
    Open Access
    C-MAT: Cross-modal aligned transformer for four-class differential diagnosis of neurodegenerative diseases using structural MRI and resting-state EEG
    (BRAC University, 2026) Ziad, Fahad Nadim; Salim, Sharika; Pal, Tonmoy; Rupom, Rubayet Hassan; Rahman, Chowdhury Mofizur; Alam, Md. Golam Rabiul; Department of Computer Science and Engineering
    Neurodegenerative diseases, including Alzheimer disease (AD), frontotemporal dementia (FTD), and Parkinson disease (PD), affect more than 57 million people globally and remain difficult to distinguish because symptoms overlap and MRI/EEG biomarkers are modality dependent. This thesis presents C-MAT (Cross-Modal Aligned Transformer), a unified multimodal framework for four-class classification of AD, FTD, PD, and healthy controls using T1-weighted structural MRI and resting state EEG. The model combines a shared pretrained image encoder, modality specific attention pooling, a dual-path EEG branch that fuses pseudo-image and raw-feature representations, and a sigmoid-gated fusion mechanism with learnable null tokens for missing modalities. C-MAT was developed through three iterative pipeline versions over nine months, with forensic debugging used to expose cross-dataset subject-ID collisions, EEG reprocessing reuse, and pretrained-weight configuration errors. The final system was evaluated on 1,102 subjects from 10 open datasets spanning seven countries, using subject-level stratified five-fold cross-validation and dataset-prefixed identifiers. ConvNeXt-Tiny achieved a macro F1 of 0.611 ± 0.029, and DeiT-Tiny achieved 0.590 ± 0.044, outperforming classical baselines such as PSD+SVM (0.309) and Random Forest (0.453). On the strictly paired 25-subject MRI+EEG subset, CMAT DeiT reached a mean macro F1 of 0.702, exceeding late-fusion baselines such as ResNet50 Fusion (0.482) and DeiT Fusion (0.296). Class-balanced focal loss improved macro F1 by +0.048 over cross-entropy. GradCAM and fusion-gate analysis showed clinically plausible patterns, including hippocampal and frontal emphasis for AD/FTD and greater EEG reliance for PD. To the best of our knowledge, CMAT is among the first MRI+EEG systems for four-class neurodegenerative disease classification evaluated under subject-level multi-site validation
  • listelement.badge.dso-type Item ,
    Open Access
    Comprehensive study on context-aware and behavioural anomaly risk assessment and prevention system : A hybrid AI-driven approach to risky driving pattern detection
    (BRAC University, 2026-04) Ahmed, S. M. Shakil; Tabassum, Nahiyan; Haque, Nusrat Zahan; Mahmud, Nusrat; Alam, Md. Golam Rabiul; Department of Computer Science and Engineering
    The number of car accidents is occurring on a global basis at an unprecedented scale, and as a consequence, this leads to an increasing demand for improving the driving risk assessment systems. Existing systems are often inadequate, as they fail to sufficiently incorporate both driver behavior and environmental context and thus often fail to perform well in real-world driving scenarios. In this work, we propose a hybrid and context-aware driving risk assessment framework, which models driver behavior and environmental conditions by employing a progressive fusion-based framework. The framework utilizes the US Accidents dataset (2016-2023) as context and risk factor, and a mobile-phone based Driving Behavior dataset capturing driving behavior. Individual classification and regression models are trained using custom built neural networks in the first stage in order to learn feature representations that are non-linear. Later, two fusion strategies are introduced, including decision level fusion (combining neural networks using probability combinations) and late fusion (combining the prediction of the classifier models with the predicted risks from context information), respectively. Besides this, in order to effectively model driver behavior over time, sequence based learning is employed to model behaviors, and a refinement layer is applied to further boost the classification accuracy of models. Our proposed approach generates a continuous and interpretable risk score and outperforms traditional methods and the single classifier by producing satisfactory classification performances. This work also emphasizes that the combined use of deep learning and fusion is an efficient approach for large-scale, reliable and multimodal driving risk assessment.
  • listelement.badge.dso-type Item ,
    Open Access
    SPRINT: Sensitivity-guided pruning for inference-time adaptation of LLMs
    (BRAC University, 2026-04) Dieyaz, Asief Iqbal; Hassan, Nahid; Yousuf, Faiyaz Bin; Hossain, Muhammad Iqbal; Department of Computer Science and Engineering
    Large Language Models have transformed artificial intelligence, showing strong performance across a wide range of natural language processing tasks; however, their computational and memory demands make deployment on resource-constrained hardware, such as personal laptops, local servers, and edge devices, largely impractical. Existing compression techniques including pruning, quantization, knowledge distillation, and Low-Rank Adaptation reduce model overhead but apply a fixed compression level at deployment time, leaving them unable to respond to changes in prompt complexity, available memory, or latency requirements during inference. This work presents SPRINT, a dynamic hybrid inference framework that selects pruning intensity on a per-prompt basis at runtime. The system consists of three components: a Learned Complexity Router (LCR) built on a fine-tuned BERT-mini backbone that predicts each prompt’s sensitivity to pruning, a Double Deep Q-Network (DDQN) controller that selects pruning actions from a 17- option discrete space using a 10-dimensional state vector combining hardware telemetry, router scores, and early backbone signals, and a structural pruning engine that physically removes transformer layers to produce real latency reductions. Experiments were conducted on Llama-2-7B using a 10,000-prompt dataset drawn equally from GSM8K, MBPP, WikiText-2, MMLU, and BoolQ. The LCR achieved a Spearman rank correlation of ρ = 0.797 (95% CI: [0.779, 0.817]) and R2 = 0.633 against oracle sensitivity labels on the held-out test set, exceeding the target threshold of ρ ≥ 0.70 across all five benchmark domains. Across 2,000 held-out test episodes, the DDQN controller reduced average inference time from 1,287.61 ms to 875.39 ms, yielding a 32.0% average speedup. Parameter count dropped from 4,714.3 MB to 3,113.5 MB, a reduction of 1,600.8 MB (34.0%), while peak VRAM consumption remained stable at approximately 4.75 GB. Total routing and action-selection overhead averaged just 17.77 ms, corresponding to approximately 2.0% of total latency. Comparison against SparseGPT, Wanda, and LLM Pruner confirmed that unstructured weight sparsity does not reliably convert to latency reduction on standard GPU hardware, whereas SPRINT’s structural layer removal yields predictable speedups without requiring sparse kernel support.
  • listelement.badge.dso-type Item ,
    Open Access
    A bilingual study of socio-cultural bias in large language models through BanglaBBQ and a post processing mitigation pipeline
    (BRAC University, 2026-04) Tasnia, Taeeba; Khan, Ashika Habib; Gomes, Sumit Anthony; Tanha, Tahseen; Saha, Provat; Hossain, Muhammad Iqbal; Anwar, Md. Tawhid; Tanvir, Sifat; Department of Computer Science and Engineering
    Large language models have achieved impressive progress in natural language understanding, but their application in practice still brings to light a vexed and understudied issue: social bias. The majority of existing bias benchmarks were constructed with largely Western, English-centric contexts, and low-resource languages and culturally diverse societies have a big gap. This gap is filled in this paper by two related contributions. We present our first bias assessment benchmark, first, the BanglaBBQ, the first bias assessment system tailored to the Bangladeshi sociocultural environment, with nine types of bias, four of them adapted to the original BBQ framework, and five newly created, such as Regional Affiliation, Educational Background, Marital Status, Mental Health, and Politics, based on recorded sociocultural realities of Bangladesh. The dataset is bilingual with structurally aligned entries in English and Bengali allowing comparison across languages. Second, we introduce SafeLLM, a threestep inference-time bias mitigation pipeline that can be trained without retraining models or having access to weights. SafeLLM uses a sensitivity layer restructuring prompts and then inferring, a bias evaluator indicating stereotype-based predictions on the sample-level and a counterfactual grounding phase that fixes identity-sensitive mistakes by exchanging the features under protection and sampling the output. Four multilingual LLMs (LLaMA- 3.1-8B, LLaMA-3.3-70B, LLaMA-4-Scout-17B, and Gemini 2.0 Flash-Lite) are tested on both languages and all categories of bias. We find a pattern of performance decreases on culturally specific templates, significant cross-lingual accuracy differences, and a model-scale dependence in the effectiveness of inference-time interventions to decrease bias. Collectively, both BanglaBBQ and SafeLLM provide a basis to culture-specific bias measurement and mitigation in multilingual AI systems.