GAFFDNet: GTE-based adaptive fusion and full-scale decoding for referring image segmentation

bracu.degree.levelPostgraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorAlam, Md. Golam Rabiul
dc.contributor.authorChowdhury, Bushra Rafia
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-02-17T04:01:47Z
dc.date.available2026-02-17T04:01:47Z
dc.date.copyright2025
dc.date.issued2025-11
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 47-49).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2025.en_US
dc.description.abstractReferring Image Segmentation (RIS) requires a precise understanding of both complex visual scenes and natural language expressions to segment target objects accurately. Despite recent advancements, traditional transformer-based methods often face challenges with ambiguous expressions, weak vision-language alignment, and limited contextual reasoning. These limitations hinder the ability to capture nuanced interactions between language and visual context, particularly in cluttered or complex scenes. This research proposed GAFFDNet, an innovative paradigm for referring image segmentation that addresses these challenges. A multi-stage contrastively trained model (GTE) is employed to produce semantically rich and discriminative textual embeddings. A dynamic multimodal fusion module is designed to adaptively integrate visual and linguistic information, which allows the network to focus on the most relevant cues. A full-scale or multi-scale feature aggregation based decoder is incorporated to facilitate dense information exchange across different scales, thereby enhancing both fine-grained spatial detail and global context understanding. Extensive evaluations conducted on three standard benchmark datasets, RefCOCO, RefCOCO+, and G-Ref demonstrate that GAFFDNet consistently outperforms existing competitive methods. These results validate the effectiveness of the proposed approach in achieving precise, context-aware segmentation aligned with complex referring expressions.en_US
dc.description.degreeMaster of Science in Computer Science and Engineering
dc.description.statementofresponsibilityBushra Rafia Chowdhury
dc.format.extent49 pages
dc.identifier.otherID 23366010
dc.identifier.urihttp://hdl.handle.net/10361/27530
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectReferring image segmentationen_US
dc.subjectImage segmentationen_US
dc.subjectPixel-level segmentationen_US
dc.subjectVision–language understandingen_US
dc.subjectMultimodal fusionen_US
dc.subjectContrastive learningen_US
dc.subject.lcshDiagnostic imaging--Digital techniques.
dc.subject.lcshImage segmentation.
dc.subject.lcshContrastive linguistics.
dc.subject.lcshVisual perception.
dc.subject.lcshOptical data processing.
dc.titleGAFFDNet: GTE-based adaptive fusion and full-scale decoding for referring image segmentationen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
23366010_CSE.pdf
Size:
2.88 MB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: