Thesis (Bachelor of Science in Computer Science and Engineering)

Permanent URI for this collectionhttps://hdl.handle.net/10361/28433

Browse

Recent Submissions

Now showing 1 - 20 of 1026
  • listelement.badge.dso-type Item ,
    Open Access
    Comprehensive study on context-aware and behavioural anomaly risk assessment and prevention system : A hybrid AI-driven approach to risky driving pattern detection
    (BRAC University, 2026-04) Ahmed, S. M. Shakil; Tabassum, Nahiyan; Haque, Nusrat Zahan; Mahmud, Nusrat; Alam, Md. Golam Rabiul; Department of Computer Science and Engineering
    The number of car accidents is occurring on a global basis at an unprecedented scale, and as a consequence, this leads to an increasing demand for improving the driving risk assessment systems. Existing systems are often inadequate, as they fail to sufficiently incorporate both driver behavior and environmental context and thus often fail to perform well in real-world driving scenarios. In this work, we propose a hybrid and context-aware driving risk assessment framework, which models driver behavior and environmental conditions by employing a progressive fusion-based framework. The framework utilizes the US Accidents dataset (2016-2023) as context and risk factor, and a mobile-phone based Driving Behavior dataset capturing driving behavior. Individual classification and regression models are trained using custom built neural networks in the first stage in order to learn feature representations that are non-linear. Later, two fusion strategies are introduced, including decision level fusion (combining neural networks using probability combinations) and late fusion (combining the prediction of the classifier models with the predicted risks from context information), respectively. Besides this, in order to effectively model driver behavior over time, sequence based learning is employed to model behaviors, and a refinement layer is applied to further boost the classification accuracy of models. Our proposed approach generates a continuous and interpretable risk score and outperforms traditional methods and the single classifier by producing satisfactory classification performances. This work also emphasizes that the combined use of deep learning and fusion is an efficient approach for large-scale, reliable and multimodal driving risk assessment.
  • listelement.badge.dso-type Item ,
    Open Access
    SPRINT: Sensitivity-guided pruning for inference-time adaptation of LLMs
    (BRAC University, 2026-04) Dieyaz, Asief Iqbal; Hassan, Nahid; Yousuf, Faiyaz Bin; Hossain, Muhammad Iqbal; Department of Computer Science and Engineering
    Large Language Models have transformed artificial intelligence, showing strong performance across a wide range of natural language processing tasks; however, their computational and memory demands make deployment on resource-constrained hardware, such as personal laptops, local servers, and edge devices, largely impractical. Existing compression techniques including pruning, quantization, knowledge distillation, and Low-Rank Adaptation reduce model overhead but apply a fixed compression level at deployment time, leaving them unable to respond to changes in prompt complexity, available memory, or latency requirements during inference. This work presents SPRINT, a dynamic hybrid inference framework that selects pruning intensity on a per-prompt basis at runtime. The system consists of three components: a Learned Complexity Router (LCR) built on a fine-tuned BERT-mini backbone that predicts each prompt’s sensitivity to pruning, a Double Deep Q-Network (DDQN) controller that selects pruning actions from a 17- option discrete space using a 10-dimensional state vector combining hardware telemetry, router scores, and early backbone signals, and a structural pruning engine that physically removes transformer layers to produce real latency reductions. Experiments were conducted on Llama-2-7B using a 10,000-prompt dataset drawn equally from GSM8K, MBPP, WikiText-2, MMLU, and BoolQ. The LCR achieved a Spearman rank correlation of ρ = 0.797 (95% CI: [0.779, 0.817]) and R2 = 0.633 against oracle sensitivity labels on the held-out test set, exceeding the target threshold of ρ ≥ 0.70 across all five benchmark domains. Across 2,000 held-out test episodes, the DDQN controller reduced average inference time from 1,287.61 ms to 875.39 ms, yielding a 32.0% average speedup. Parameter count dropped from 4,714.3 MB to 3,113.5 MB, a reduction of 1,600.8 MB (34.0%), while peak VRAM consumption remained stable at approximately 4.75 GB. Total routing and action-selection overhead averaged just 17.77 ms, corresponding to approximately 2.0% of total latency. Comparison against SparseGPT, Wanda, and LLM Pruner confirmed that unstructured weight sparsity does not reliably convert to latency reduction on standard GPU hardware, whereas SPRINT’s structural layer removal yields predictable speedups without requiring sparse kernel support.
  • listelement.badge.dso-type Item ,
    Open Access
    A bilingual study of socio-cultural bias in large language models through BanglaBBQ and a post processing mitigation pipeline
    (BRAC University, 2026-04) Tasnia, Taeeba; Khan, Ashika Habib; Gomes, Sumit Anthony; Tanha, Tahseen; Saha, Provat; Hossain, Muhammad Iqbal; Anwar, Md. Tawhid; Tanvir, Sifat; Department of Computer Science and Engineering
    Large language models have achieved impressive progress in natural language understanding, but their application in practice still brings to light a vexed and understudied issue: social bias. The majority of existing bias benchmarks were constructed with largely Western, English-centric contexts, and low-resource languages and culturally diverse societies have a big gap. This gap is filled in this paper by two related contributions. We present our first bias assessment benchmark, first, the BanglaBBQ, the first bias assessment system tailored to the Bangladeshi sociocultural environment, with nine types of bias, four of them adapted to the original BBQ framework, and five newly created, such as Regional Affiliation, Educational Background, Marital Status, Mental Health, and Politics, based on recorded sociocultural realities of Bangladesh. The dataset is bilingual with structurally aligned entries in English and Bengali allowing comparison across languages. Second, we introduce SafeLLM, a threestep inference-time bias mitigation pipeline that can be trained without retraining models or having access to weights. SafeLLM uses a sensitivity layer restructuring prompts and then inferring, a bias evaluator indicating stereotype-based predictions on the sample-level and a counterfactual grounding phase that fixes identity-sensitive mistakes by exchanging the features under protection and sampling the output. Four multilingual LLMs (LLaMA- 3.1-8B, LLaMA-3.3-70B, LLaMA-4-Scout-17B, and Gemini 2.0 Flash-Lite) are tested on both languages and all categories of bias. We find a pattern of performance decreases on culturally specific templates, significant cross-lingual accuracy differences, and a model-scale dependence in the effectiveness of inference-time interventions to decrease bias. Collectively, both BanglaBBQ and SafeLLM provide a basis to culture-specific bias measurement and mitigation in multilingual AI systems.
  • listelement.badge.dso-type Item ,
    Open Access
    Detection of food adulteration using machine learning tools
    (BRAC University, 2026-01) Alamgir, Anika; Anjum, Nashita; Hossain, A. K. M. Shahadat; Deb Nath, Debpriyo Hrishikes; Rahman, Chowdhury Mofizur; Rahman, Rafeed; Department of Computer Science and Engineering
    Adulteration of food has been an urgent concern for the health and food safety of the people, and the affected area is particularly the developing countries due to the inefficient infrastructure and limited resources available to test the food items. Some traditional forms of detection, like chemical analysis and chromatographic analysis, are usually slow, costly, and require competent manpower. This paper explores an affordable and scalable Machine Learning (ML) system to identify food adulteration in sugar and poppy seed, and also another criteria is to identify the freshness of meat. The methodology of the study uses a modular approach to using self created RGB image dataset of sugar mixed with Magnesium Sulfate as adulterant which is one of its first kind in food adulteration detection of sugar and another one is of poppy seed dataset which is mixed with semolina as adulterant, collected through online to predict different fixed adulteration levels. Additionally, the dataset of meat is collected from an existing paper for further working on meat freshness detection. A self proposed ensembled model using pretrained models including: Densenet121, Conv-Next Tiny and Resnet50 with soft voting is used for detecting adulterants of sugar with Grad-CAM visualization, also showed some success in detecting other unknown adulterants. The ensembled model for sugar has shown about 88.96% accuracy. Whereas for poppyseed, Convolutional Neural Networks (CNN), and its pretrained model- EfficientNetB0 shows very high accuracy - 98.25% itself to measure the level of adulterants in class-wise manner without requiring any ensemble learning. As per previous works, there are few models that have satisfactory accuracy rate in terms of adulterant detection but there still exist issues, especially in covering diverse food items like sugar and poppyseed and as of for meat freshness, the Misclassification Cost based on some shown pretrained models was much higher in a previous work, which has been addressed by this study. For meat dataset, a Swin Transformer integrated with Explainable AI succesfully identified spots of spoilage, solving the black-box problem, showed increased accuracy upto 98% which is higher compared to that of the models shown in the existing paper from where the meat data set was collected and along with this, the study addressed the misclassification cost, an evalutation metric proposed by the original paper, which was lowered by many folds compared to the exisiting misclassification cost of the original paper. This research highlights that image data by mobile or obtained on accessible platform with AI can facilitate fast, dependable food safety testing and aid food quality inspection in supply chains.
  • listelement.badge.dso-type Item ,
    Open Access
    Integrating mental-RoBERTa and generative LLMs for phenotyping secondary insomnia: A multi-dimensional framework for etiology discovery
    (BRAC University, 2026-01) Sakib, Sadman; Hossain, Nahmar; Chowdhury, Kumar Prothom Pranto Sarma; Hakim, Shehzad; Islam, A Q M Mujahidul; Anwar, Md. Tawhid; Ahmed, Md. Sabbir; Alam, Md. Ashraful; Department of Computer Science and Engineering
    Insomnia affects a significant portion of the global population, yet distinguishing between primary and secondary insomnia—where sleep disturbances stem from underlying medical, psychiatric, or environmental factors—remains critically underexplored in computational health research. This study presents a multi-stage computational framework for identifying and phenotyping secondary insomnia from naturalistic Reddit discussions. A clinically validated dataset of 600 posts was developed, with linguistic validation confirming that secondary insomnia posts contain significantly more biomedical terminology than primary insomnia posts. Mental- RoBERTa, a domain-adapted transformer, was employed (after the initial integration of RoBERTa-Base) for identifying cases, which outperformed both traditional baselines and large language models. The trained model identified over 3,400 high confidence secondary insomnia cases from 5,000 unlabeled posts. A hybrid pipeline combining Llama-3 clinical summarization with BERTopic unsupervised clustering discovered 10 distinct etiological categories, revealing gastroesophageal reflux disease, menopausal factors, and benzodiazepine withdrawal as predominant drivers. Zero-shot emotion profiling revealed distinct psychological phenotypes across different triggers, with benzodiazepine withdrawal exhibiting the highest perplexity and environmental factors showing the greatest frustration. This framework enables the first large-scale computational characterization of secondary insomnia drivers and their psychological burdens, providing actionable insights for precision intervention design and population health surveillance in digital mental health systems.
  • listelement.badge.dso-type Item ,
    Open Access
    AI powered hospital management system (AI-HMS)
    (BRAC University, 2026-02) Nodi, Nusrat Nowshin; Khatun, Mst Sushmita; Priti, Annesha Das; Azmain, Md. Aquib; Department of Computer Science and Engineering
    The increasing complexity of modern healthcare services has highlighted the limitations of traditional Hospital Management Systems, which are primarily focused on the automation of hospital administration and electronic management of patient records, with limited support for decision-making. This thesis proposes the design and development of AI-HMS, an Artificial Intelligence-based Hospital Management System that incorporates the integration of administration and decision making through the use of artificial intelligence. The proposed system is built using a full-stack development framework, with the React.js library for the front-end, the Flask framework for the back-end, and PostgreSQL for the management of both structured and semi-structured data. Supervised learning algorithms are used for the prediction of the risk level of patients and the likelihood of hospital readmission. A Large Language Model-based artificial intelligence assistant is incorporated for the assistance of healthcare professionals and patients. Therefore, AI-HMS is based on the human-in-the-loop philosophy of artificial intelligence, where the artificial intelligence component is used as a decision-support tool. The proposed system was evaluated using synthetic data sets for healthcare services and was found to perform reliably. The system was also deployed as a web-based application, validating the proposed system. Overall, this research demonstrates the potential of using artificial intelligence for the management of hospital services.
  • listelement.badge.dso-type Item ,
    Open Access
    An algorithmic method for warehouse order picking with congestion and one-way aisle constraints
    (BRAC University, 2026-01) Bidhu, Shad Nur Mim; Arif, Abdullah; Abony, Maharin Sharif; Ahrar, Nazib Uddin; Shafin, Syed Hasibur Rahman; Chakrabarty, Amitabha; Department of Computer Science and Engineering
    Order picking is a fundamental management process in warehouse management systems and it has a great impact on the cost of operations, efficiency and quality of the services. Although the use of traditional routing and shortest path algorithms is a frequent solution to the problem of order-picking, it frequently ignores real-world constraints like the congestion and one-way aisle patterns, which can significantly impact the viability and effectiveness of routes in the real-world warehouse setting. To overcome these drawbacks, this thesis is concerned with creating and analyzing congestion-aware routing plans in a realistic warehouse environment with order-picking. A simulation based benchmarking model is suggested to compare systematically classical, heuristic, hybrid, and metaheuristic routing algorithms based on unified warehouse layouts with one-way aisle constraints and a travel cost that is a congestion cost. The structure makes it possible to model route execution, traffic congestion, and collision impacts that allow fair and consistent comparison between algorithm methods. Performance variables are calculated on standardized measures, including quality of solutions, planning efficiency, collision impact, resource use and scalability to problems of different sizes. The comparison analysis shows that even though exact routing algorithms only work well in small scale problems, heuristic routing methods are quicker but more vulnerable to congestion impacts. Hybrid routing methods are more robust since they trade off between computational efficiency and quality of solutions in congestion aware situations. For instance, Hybrid NN2opt achieved a 74.1% optimization rate with a mean planning time of only 7.68 ms, nearly matching the optimal Held-Karp baseline while using 89% less memory. In multi-robot simulations, it reduced collisions by up to 33.33% in congested narrow aisles and up to 100% in wide aisles and maintained more stable performance as robot density increased from 3–5 to 10–15 robots. These results underscore the significance of the inclusion of realistic operation constraints in the modeling of warehouse routing and offer a feasible base on how to enhance the efficiency of order-picking in a big and dynamic warehouse setup.
  • listelement.badge.dso-type Item ,
    Open Access
    AI-enhanced vulnerability detection: A machine learning approach to automated penetration testing
    (BRAC University, 2025-10) Maruf, Mahmud Mostofa Al; Tabassum, Mahmuda; Disha, Radiah Reaz; Emon, Salauddin Ahmed; Hossain, Muhammad Iqbal; Department of Computer Science and Engineering
    Cybersecurity is an issue that is being compounded by the fast development of digital infrastructure as cyberattacks grow more sophisticated and frequent. Vulnerability detection and penetration testing are some of the most important challenges in the area of cybersecurity as they allow organizations to highlight the security vulnerabilities prior to the exploitation. The existing traditional penetration testing is, however, largely manual, expensive, inefficient, and time-consuming, especially in large scale dynamic networks. This paper presents an improved vulnerability detection system using AI that can be used to automate penetration testing to enhance the ability of the system to detect threats more effectively. It incorporates automated attack path exploration, AI based threat intelligence and adversarial defense to increase the precision, versatility and dependability of penetration tests. To make security assessment transient and interpretable, explainable-AI methods are also included. Findings show that AI-based penetration testing also saves considerable time in human work, and it detects high-risk threats more efficiently through the analysis of large datasets, including complicated attack patterns. This piece is emphasizing a change to the conventional way of approaching cybersecurity as it provides a more effective and scalable alternative to the conventional way of doing it.
  • listelement.badge.dso-type Item ,
    Open Access
    Consumer and business owner rights and legal access: An HCI based approach
    (BRAC University, 2026-01) Meherin, Tasmiah; Khan, Tanjidul Alam; Kareeb, Khan Mohammad Al; Azad, Shahriar; Tasnim, Obyda; Mukta, Jannatun Noor; Department of Computer Science and Engineering
    Nowadays, people are being cheated every day while buying groceries. The sad thing is that even if people take measures to avoid awareness, the results are revealed very late. Another issue is that people still do not fully know what laws apply and what steps should be taken if they are deceived while buying a product. And because of this, the syndicate system in the market is increasing day by day. The main goal of our research is that if people are victims of fraud when buying a product in the current market, they can easily access the laws that exist for that case which are included in consumer act. The types of cases that fall under this category are expired products, products sold at a price higher than the original price, adulterated and counterfeit products, etc. We train the models of this earning model including Naive Bayes, Decision Tree, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), LDA, Linear Regression, Logistic Regression, Decision Stump with data collected through the survey (n=1500) according to the type of fraud. After analyzing the models, it was seen that SVM gives the most effective results which showed accuracy of 98%, recall (97%), F1-score (97%), precision (97%). This study will help the people of Bangladesh to easily know which law is applicable for that fraud and what measures need to be taken if they are a victim of fraud when buying any product in the local market. By doing this, it will ultimately create awareness among the businessmen in the country, protect the public health of the common people, and eliminate the syndicate system in the local market. This will improve the economic system of Bangladesh on the one hand as well as the development of the country.
  • listelement.badge.dso-type Item ,
    Open Access
    A multimodal AI-based agricultural assistance system: Enhancing farmer decision-making through visual question answering (VQA) in Bangladesh
    (BRAC University, 2026-01) Arjan, Promit Dey Sarker; Mati, Mrittika Devi; Basak, Argha; Bari, Saib Sadman; Chakrabarty, Amitabha; Bhoumik, Partha; Department of Computer Science and Engineering
    Agriculture plays a vital role in Bangladesh, employing approximately 43% of the population, yet many farmers lack formal education and expert guidance, leading to uninformed decision-making. This research explores the development of a multimodal Visual Question Answering (VQA)-based agricultural assistance system that integrates computer vision, natural language processing (NLP), and speech processing to provide real-time, voice-based responses in Bangla, reducing the immediate need for agricultural experts. Instead of relying solely on CNN-based classification, the system employs Vision–Language Models (VLMs) with visual encoders for crop disease understanding, a Large Language Model (LLM) for agricultural query answering, and speech models for Bangla voice interaction. The system offers an intuitive, offline-compatible mobile solution tailored for farmers in low-resource environments. A conversational agricultural diagnostic dataset was developed containing 96,003 expert-style diagnostic entries with 9,138 rice leaf images across six classes. A structured dataset covering Bangladesh’s diverse crops, pests, and diseases ensures AI model optimization and practical usability. A comparative evaluation assesses the multimodal AI system (Image + Text + Voice) against traditional single-input models (Image-only, Text-only) based on diagnostic accuracy, response relevance, and user satisfaction. In the final offline pipeline, MobileNetV3-Small achieved 95.35% disease classification accuracy with stable mobile CPU latency ( 12– 18 ms), Gemma 3–1B produced the strongest overall response quality among tested lightweight LLMs (METEOR = 0.1095, cosine similarity = 0.539, and BERTScore F1 ￿ 0.70, BertScore Precision: 0.72, BertScore Recall: 0.69), and Whisper-Base enabled faster-than-real-time Bangla speech recognition (RTF = 0.7859, WER = 0.353). This approach minimizes dependence on agricultural experts, saving time and costs while positively impacting the economy. A farmer field survey (n = 30) using five evaluation parameters indicates strong real-world usability: 92% rated the system helpful/very helpful, 88% found it easy to use, 84% reported successful offline usage, 86% expressed satisfaction with response clarity, and 82% reported improved confidence in decision-making. Ultimately, this research seeks to bridge the digital divide in agriculture, empowering farmers with an AI-driven solution for real-time problem-solving and enhanced food security.
  • listelement.badge.dso-type Item ,
    Open Access
    A hybrid deep learning framework for multi-class ICH detection and severity assessment from CT scans
    (BRAC University, 2026-02) Morshed, Atik; Tamanna, Rifah Jahan; Alam, Fareha; Lamia, Ishrat Jahan; Danial, A.B.M; Alam, Md. Ashraful; Department of Computer Science and Engineering
    Intracranial hemorrhage (ICH) is a critical neurological condition that requires rapid and accurate assessment for effective clinical management. Analysis of head CT scans using automated systems can be used to help clinicians as it enhances the reliability of detections and can be used to prioritize in the emergency room. Nevertheless, the current approaches tend to detect hemorrhage only and use considerable annotated datasets, whereas the severity assessment is still underresearched because of the lack of ground-truth labels. In this thesis, we will present a hybrid deep learning model to detect multi-class intracranial hemorrhage and severity stratification by proxy as an element of clinical metadata and head CT images. The framework utilizes an EfficientNet-B0 backbone, to obtain slice level visual features, which are pooled in order to obtain low contextual information between slices. The feature fusion is used to add more contextual cues by incorporating patient level metadata. The model is trained under a multi-task learning environment, where it concurrently detects the subtypes of hemorrhage and classifies the category of the severity. Since the CQ500 dataset lacks explicit severity annotations, proxy hemorrhage volume estimates and dataset specific quantile thresholds are used to create severity labels in a weakly supervised fashion. This allows relative stratification of severity in the dataset as opposed to clinical grading of severity. Transfer learning is used to pretrain on the RSNA Intracranial Hemorrhage dataset and then fine-tune on CQ500 to enhance generalization with limited data. Experimental results demonstrate competitive patient-level detection performance across multiple hemorrhage subtypes, achieving high ROC-AUC values on CQ500. The severity classification task achieves moderate and consistent performance, with most errors occurring between adjacent severity levels, reflecting the continuous nature of hemorrhage extent and the proxy-based labeling scheme. Overall, this work demonstrates the feasibility of combining multimodal data, weak supervision, and multi-task learning for intracranial hemorrhage analysis under constrained annotation settings. While not intended for direct clinical deployment, the proposed framework provides a foundation for future research incorporating expert-annotated severity labels and external validation.
  • listelement.badge.dso-type Item ,
    Open Access
    Enhancing knee osteoarthritis diagnosis with AI: A deep learning perspective
    (BRAC University, 2026-01) Saha, Arnab; Mahbub, Mayeesha; Naem, Jannatul; Rahman. Adiba; Mobashir, Mohammed Munif; Karim, Dewan Ziaul; Department of Computer Science and Engineering
    Knee Osteoarthritis is one of the most concerning diseases in the current world. A lot of people are suffering from Knee Osteoarthritis which is a common disease now-a-days. In the current world, the number of knee patients is increasing day by day. The risk of KOA increases for anyone who gets older or carries excess weights of body. Stress also causes KOA. Additionally genetic problems or joint injuries can cause KOA. Knee Osteoarthritis are two types. In the primary stage(KL-1 and KL-2) it is called Osteopenia and in the second stage(KL-3 and KL-4) is called Osteoporosis are significant global health concerns, leading to increased bone fragility and fracture risk. Usually patients feel pain in and around the knee, pain is frequently described as a stabbing, sharp, or dull ache. Effective clinical intervention depends on early and precise detection using X-ray imaging. However, manual diagnosis is difficult and subject to inter-observer variability due to minute differences in bone density and texture. In this paper, a new, lightweight architecture for a convolutional neural network is proposed, termed Res-SE LiteNet, designed to classify knee radiograph images into three classes: Normal, Osteopenia, and Osteoporosis. In order to minimize computational parameters, thereby making the model practical for a clinical environment, the proposed architecture utilizes Depthwise Separable Convolutions. To enhance feature sensitivity, we implemented a Channel-wise Attention Mechanism via Squeeze-and-Excitation (SE) blocks and Residual Skip Connections to prevent the degradation of low-level spatial details. A Hybrid Global Pooling strategy, combining Global Average and Max Pooling, was employed to capture both global bone density statistics and focal pathological markers. A dataset publicly available in Mendeley was used to train and validate the proposed model. The model was able to attain a peak validation accuracy of 90.53%. Compared to conventional models such as the VGG16 or ResNet50 models, the model is able to maintain high accuracy while using a much smaller number of parameters. By comparing computationally complex pre-trained models, our proposed lightweight model achieves comparable diagnostic accuracy with a significantly reduced parameter footprint. As a result the model can facilitate high-speed processing without compromising model accuracy.
  • listelement.badge.dso-type Item ,
    Open Access
    Structural variant analysis in human genome using transformer based genomic language model
    (BRAC University, 2026-01) Islam, Shaker; Sami, Md. Ashraful Islam; Rohan, Amin Mohammad; Deb, Jhishan; Alam, Md. Golam Rabiul; Shatabda, Swakkhar; Department of Computer Science and Engineering
    Long read sequencing is a DNA sequencing technique that makes it possible to sequence long DNA fragments, offering an unprecedented opportunity to resolve structural variants (SVs) in genomes. Structural variants (SVs) are changes in DNA that play an important role in genome diversity, evolution, and disease. Detecting SVs in the human genome is challenging because of differences in genome structure, complexity, and limited labeled data. Existing long read structural variant detection methods depend on predefined rules or heuristic strategies, which do not fully capture the intricate nature of SV signatures. To overcome these limitations, we introduce a transformer driven model for analyzing structural variants in the human genome called SVBERT. Our approach first extracts SV signatures from sequence alignments and assembles local regions to generate paired sequence inputs. Local alignment features are processed by a modified convolutional neural network (CNN) encoder, while BERT generates context aware sequence embeddings. A fusion module then combines these features using cross attention followed by a transformer encoder. Finally, specialized prediction heads perform classification, breakpoint regression, genotype calling, and confidence scoring. Post processing with confidence based filtering produces high quality structural variant calls. Validation across human genome sequencing datasets shows improved detection of various SV types. These results demonstrate the strong potential of transformer based genomic language models for advancing SV analysis in both research and practical applications.
  • listelement.badge.dso-type Item ,
    Open Access
    Deep learning based Braille character to Bangla voice conversion system
    (BRAC University, 2026-01) Chowdhury, Ariq Sadiq; Bakhtiar, Rafid Bin; Chakraborty, Aritra; Foraejy, Aowfi Adon; Chakrabarty, Amitabha; Department of Computer Science and Engineering
    Braille is an essential mode of communication for visually impaired individuals, enabling them to read and understand contexts through a tactile system of raised dots. However, certain regions worldwide, especially Bangladesh, have limited access to digital solutions for Braille conversion, which poses a challenge for those individuals. This thesis presents a method for converting Braille Characters to Bangla Voice through a deep-learning based system and thus enhancing accessibility for visually impaired individuals as well as general individuals who cannot understand Braille. In this research we have addressed the scarcity of high volume labeled data consisting of more than 10000 high-resolution labeled data with approximately 1 million labeled braille instances by strictly maintaining Library of Congress spatial standards. This research simulataniously performs two different architecture one being a multi-stage segmentation-classification and other one being YOLO based unified architecture. The unified architecture vastly outperforms the traditional method by achieving Mean Average Precision (mAP@50) of 98.5 percent in character extraction. The recognized text is then converted into natural-sounding Bangla speech using a TTS engine. This system is evaluated through accuracy, processing speed, and friendly user experience, demonstrating its potential to bridge the communication gap for the visually impaired in Bangla-speaking communities.
  • listelement.badge.dso-type Item ,
    Open Access
    A robust federated learning privacy framework: Integrating multi-party computation and advanced privacy-preserving techniques for secure data collaboration
    (BRAC University, 2026) Tasmim, Ahana; Musarrat, Aneela; Rahman, Md. Naim; Khan, Md. Sakib; Asaduzzaman, Md.; Mostakim, Moin; Department of Computer Science and Engineering
    Federated Learning(FL) can train models without exposing raw client data, but still vulnerable to privacy leakage and poisoning attacks. Present methods like secure aggregation tend to offer little to no malicious update detection. Other research methods to mitigate this problem also fail to give a proper balance between privacy and robustness, or often give an impractical premise of needing multiple servers, creating a complicated application and deployment. This paper proposes DDFed-Markov, a probabilistic dual-defense federated learning framework using markovian chain that simultaneously improves privacy without compromising robustness. The framework uses a combination of noise injection using Markovian probability, Fully Homomorphic Encryption (FHE), and a feedback-driven strategy based on consensus with the collaboration of the Dual Defense Federated Learning Framework. While other models that use the concept of noise work in a similar way, the Markovian noise remains different because of its state-dependent nature. It introduces temporal correlation that increases resistance to adversaries while maintaining learning efficiency. Encrypted updates are aggregated using CKKS homomorphic encryption, allowing similarity-based aggregation without the need for multiple servers, preserving FLs’ hub and spoke topology. Also, client-side consensus is applied to filter malicious updates using adaptive thresholds. Experimental evaluations on publicly available data sets like MNIST have shown that DDFed-Markov has managed to achieve a comparative accuracy while effectively prevents potential model poisoning and successfully preserving privacy compared to existing frameworks.
  • listelement.badge.dso-type Item ,
    Open Access
    Exploring cross-domain Bangla text summarization using large language models
    (BRAC University, 2026-01) Kotha, Eshika Ebnat; Chowdhury, Abtahi Bin Jahangir; Anan, Rafiyad Khan; Showkat, Subha Naj; Afridi, Sayed; Sadeque, Farig Yousuf; Siddiqui, Md. Saiful Bari; Department of Computer Science and Engineering
    Automatic text summarization is a critical tool for managing the growing volume of digital content, yet effective summarization remains challenging for low-resource languages such as Bangla. This thesis investigates the capability of large language models (LLMs) to perform cross-domain Bangla text summarization under a strictly zero-shot setting. Rather than proposing a new summarization model, the study focuses on a systematic and reliable evaluation of existing models across heterogeneous domains. Summarization outputs are generated from two distinct datasets: a real-world Bangla news corpus (Prothom Alo) and the benchmark XL-Sum (Bangla) dataset. A diverse set of encoder–decoder and decoder-only LLMs is evaluated using a multi-layered assessment framework that combines traditional automatic metrics, blind LLM-as-a-Judge evaluation, SBERT-based semantic similarity analysis, and an automated error taxonomy. We assume that we need a more robust comparison beyond surface level lexical matching, which is found ineffective for Bangla abstractive summarization. However, our experimental results show that those lexical metrics (such as ROUGE and BLEU) are generally insufficient to reflect semantic quality for Bangla summary since near zero scores (‘0’scores) appear in the case of coherent summarization. In contrast, the semantic analysis shows that Banglaspecific encoder-decoder models including BanglaT5 and mT5 significantly better perform than both the multilingual and decoder-only in domains. Decoder-only models are observed to behave erratically and incline towards either ungrammatical extraction or hallucination, as is systematically verified using semantic similarity patterns and error taxonomy analysis. The findings show that fine summarization in Bangla is insensitive to surface fluency or lexical overlap but dependents on semantic abstraction and faithfulness. We believe that by presenting an exhaustive and behavior-aware evaluation framework, we are able to give practical advice for the future Bangla summarization work so as to demonstrate the importance of language-wise evaluation methodologies especially for low resource languages.
  • listelement.badge.dso-type Item ,
    Open Access
    A computer vision driven ecosystem for cattle monitoring: Multi-disease classification with severity grading, multi-view individual identification, and weight estimation
    (BRAC University, 2026) Raj, Aabu Yousuf; Hasan, Md. Rakibul; Arafat, Abdur Rahman; Rahman, S.M. Sadman; Jalal, Md. Roman Bin; Mukta, Jannatun Noor; Department of Computer Science and Engineering
    Effective cattle monitoring and management are essential for sustainable dairy and beef production in Bangladesh, where constant veterinary care and extensive monitoring are economically challenging. This thesis introduces a computer vision based cattle monitoring ecosystem that create a (i) multi-disease classification and severity grading, (ii) multi-view unique cattle identification, and (iii) body weight estimation with the use of four view RGB images, based on a custom-collected, curated dataset with publicly available sources. To analyze the disease, a hierarchical deep learning architecture is suggested to categorize Lumpy Skin Disease (LSD), Foot-and-Mouth Disease (FMD), Infectious Bovine Keratoconjunctivitis (IBK), and Healthy cattle along with grading the diseased cattle as Stage-1 (mild), Stage-2 (moderate), and Stage-3 (severe) cases. The severity grading is performed based on a clinically validated symptom profile and is supervised and confirmed by a District Livestock Officer (veterinarian). The cross-attentional multi-task architecture has the highest accuracy score of 89.96% in disease classification, 83.75% in severity staging, and 85.88% in hierarchical setup (predicts disease and then predicts severity). In the case of individual cattle recognition and weight estimation, a multi-view appearance-based recognition system is constructed based on left, right, front, and back views of cattle images where at first YOLOv8s is trained to extract the exact cattle region which have an IoU of 0.93, map@0.5 0.97. The accuracy of the identification system is determined at Rank-1: 96.40% and 74.56% while Rank-5: 100% and 89.80% for ConvNext-Tiny on two different protocols of Leave One View Out testing method and Cross-view-angle respectively which has been done to evaluate cattle identification from different angles or possess and consider the case of cattle with very similar patterns and colors. To estimate body weight, a regression model is used with multi-view RGB images and metadata that includes id, sex, breed, age and live weight where the best single-model baseline is the DN121-tuned regressor, achieving an MAE of 36.99 KG which improves to an MAE of 35.10 KG with ensemble strategy. In general, the suggested system proves that even without the use of invasive sensors, an RGB-based, low shot, non-invasive, and affordable vision-based system can be used successfully in the diagnosis of diseases and their severity grading, the unique identification of individuals among cattle, and the estimation of their live weight.
  • listelement.badge.dso-type Item ,
    Open Access
    Evaluating the efficacy of incentives in enhancing voluntary blood donation: A behavioral and operational approach in Bangladesh
    (BRAC University, 2026-01) Mysha, Sherajum; Shishir, Md. Mehedi Hasan; Alif, Mehedi Hassan; Chowdhury, Farida; Ahmed, Md. Sabbir; Department of Computer Science and Engineering
    Blood donation is a vital component of public health, yet many regions, including Bangladesh, face challenges with insufficient voluntary donations due to various practical and motivational barriers. This thesis investigates a hybrid model that retains the humanitarian essence of donation while introducing strategic incentives, such as transportation reimbursements and potential e-commerce benefits, to overcome logistical and financial deterrents. To validate this framework, the study involved the development of a functional mobile application, Shebok, designed to transparently manage these rewards and streamline the donor-recipient matching process. The research further explores the operational efficiency of requiring donors to submit recent donation history, aiming to reduce the time spent on health screenings and ease the burden on medical staff. Based on qualitative data gathered from (N = 53) participants (50 general users and 3 institutional stakeholders) via semistructured interviews, findings confirm that non-monetary incentives are an effective mechanism to significantly increase donor motivation without compromising ethical standards. Ultimately, this research outlines a sustainable, donor-friendly system that balances humanitarian values with practical incentives to improve the overall effectiveness of blood donation programs in Bangladesh.
  • listelement.badge.dso-type Item ,
    Open Access
    Uncertainty-aware federated dual-branch multimodal fusion for privacy-preserving skin lesion diagnosis
    (BRAC University, 2025-06) Roup, Ryan Azim; Sultana, Sabrina; Tahsin, Saba; Morshed, Mostakim; Khanam, Marzia; Siddiqui, Md. Saiful Bari; Sadeque, Farig Yousuf; Department of Computer Science and Engineering
    Skin lesion diagnosis is well-suited to data-driven methods, yet its real-world use is often hindered by strict privacy rules, fragmented data ownership, and inconsistent imaging conditions. To overcome these barriers, our work introduces a robust privacy-preserving framework that combines multimodal representation learning with decentralized training which reduces the need to centralize sensitive clinical data. We propose a unified dual-branch framework that supports (i) task-level multitask learning by coupling lesion classification with lesion segmentation using a shared Swin Transformer encoder and an Attention U-Net decoder, and (ii) modality-level learning by fusing image features with auxiliary patient metadata through configurable fusion strategies. To address client heterogeneity under federated learning, we introduce a Dynamic Uncertainty-Aware Divide2Conquer (DUA-D2C) aggregation strategy that adaptively weights client updates based on predictive performance and uncertainty, improving robustness in non-IID settings. Extensive evaluations on HAM10000 and MILK10K demonstrate that multimodal fusion improves over image-only baselines, and that DUA-D2C preserves or improves performance on HAM10000 while constraining degradation on the more heterogeneous MILK10K setting.
  • listelement.badge.dso-type Item ,
    Open Access
    Regional clustering, time series forecasting, and satellite-to-ground data mapping for environmental analysis in Bangladesh using explainable AI
    (BRAC University, 2026-02-06) Tamanna, Thufa Anwar; Joy, Shah Newaz Khan; Masum, Zayed; Abdullah, Fariha Binte; Anam, S M Kafi; Rahman, Chowdhury Mofizur; Karim, Dewan Ziaul; Department of Computer Science and Engineering
    Air pollution and changes in the weather make it hard to keep an eye on the environment and protect public health, especially in areas where there aren’t many ground-based sensors or they are spread out unevenly. Ground measurements give accurate local readings, but they fail to encompass much of ground data. Spatial coverage limits analysis on a regional scale. Satellite observations provide extensive spatial data but lack local accuracy. This thesis suggests a complete multimodal machine learning framework that combines satellite-derived spatial features, ground-based environmental measurements, regional clustering, and time-series forecasting to better predict and understand air quality and important weather variables. The suggested framework uses gradient-boosted regression models for multimodal learning and pre-trained convolutional neural networks (ResNet-18, ResNet-50, and DenseNet- 121) to get features from satellite images. Time-series forecasting is used to find patterns over time and changes in environmental factors that happen at different times of the year. Unsupervised regional clustering is used to find groups of districts that have similar pollution and weather patterns. The framework is assessed against various air pollutants (SO, NO2, O3, PM2.5, PM10) and meteorological factors (solar radiation, relative humidity, and rainfall) utilizing standard performance metrics, such as RMSE, MAE, and R2. The experimental results demonstrate the effectiveness of multi-modal fusion is highly dependent on the target and region. For variables that were affected, like sulfur dioxide and rainfall, explained variance went up a lot and prediction error went down a lot. On the other hand, O3 and solar radiation got a little better. Again, NO2, particulate matter, and relative humidity were best modeled using only ground-based data. This suggests that there are many localized emission sources and processes that happened close to the ground. To improve the transparency of the model, explainable artificial intelligence (XAI) methods were added through global feature importance analysis with XGBoost. This gave us a better idea of what environmental factors were most important for making predictions over time. The results underscore the significance of pollutant- and region-specific modeling approaches and illustrate that multimodal learning can enhance environmental prediction when informed by the physical and statistical attributes of the target variables.