Alam, Md. Golam RabiulChowdhury, Bushra Rafia2026-02-172026-02-1720252025-11ID 23366010http://hdl.handle.net/10361/27530Cataloged from PDF version of thesis.Includes bibliographical references (pages 47-49).This thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2025.Referring Image Segmentation (RIS) requires a precise understanding of both complex visual scenes and natural language expressions to segment target objects accurately. Despite recent advancements, traditional transformer-based methods often face challenges with ambiguous expressions, weak vision-language alignment, and limited contextual reasoning. These limitations hinder the ability to capture nuanced interactions between language and visual context, particularly in cluttered or complex scenes. This research proposed GAFFDNet, an innovative paradigm for referring image segmentation that addresses these challenges. A multi-stage contrastively trained model (GTE) is employed to produce semantically rich and discriminative textual embeddings. A dynamic multimodal fusion module is designed to adaptively integrate visual and linguistic information, which allows the network to focus on the most relevant cues. A full-scale or multi-scale feature aggregation based decoder is incorporated to facilitate dense information exchange across different scales, thereby enhancing both fine-grained spatial detail and global context understanding. Extensive evaluations conducted on three standard benchmark datasets, RefCOCO, RefCOCO+, and G-Ref demonstrate that GAFFDNet consistently outperforms existing competitive methods. These results validate the effectiveness of the proposed approach in achieving precise, context-aware segmentation aligned with complex referring expressions.49 pagesenBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.Referring image segmentationImage segmentationPixel-level segmentationVision–language understandingMultimodal fusionContrastive learningDiagnostic imaging--Digital techniques.Image segmentation.Contrastive linguistics.Visual perception.Optical data processing.GAFFDNet: GTE-based adaptive fusion and full-scale decoding for referring image segmentationThesis