OmniNet: a hybrid deep learning framework for robust semantic segmentation

bracu.degree.levelPostgraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorAlam, Md. Golam Rabiul
dc.contributor.authorIfty, Fahim Ahmed
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-03-04T04:10:14Z
dc.date.available2026-03-04T04:10:14Z
dc.date.copyright2025
dc.date.issued2025-10
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 45-47).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science, 2025.en_US
dc.description.abstractSemantic segmentation empowers computers to interpret visual scenes in a structured and meaningful manner, accurately delineating the precise boundaries of every object in an image at the pixel level. However, existing CNN-based models struggle with long-range context, transformer-based approaches are computationally heavy and often miss local detail, and many hybrid designs neglect explicit multi-scale features or refined attention, while also suffering from class imbalance and noisy annotations. This research proposes an encoder–decoder architecture that combines a ResNet152 backbone, Atrous Spatial Pyramid Pooling (ASPP), and a Transformer Refine Block, leveraging the strengths of convolutional features and lightweight self-attention to capture both multi-scale semantics and long-range dependencies. In the decoder, a UNet-style upsampling path with skip connections and CBAM attention preserves spatial detail and refines boundaries, yielding coherent pixel-level predictions across scales. The training pipeline applies label cleaning and robust augmentation (MixUp, CutMix) and optimizes a composite loss (Lovasz–Softmax, Dice Loss, Boundary Loss) to counter label noise and the long-tail distribution. Evaluated on the CamVid benchmark, our method achieves 86.53% mean IoU and 95.99% pixel accuracy, outperforming recent efficient residual attention networks while prioritizing segmentation quality over inference speed. These results demonstrate that combining multi-scale convolutional features with lightweight self-attention and boundary-aware optimization delivers state-of-the-art accuracy under noisy labels and severe class imbalance in urban scene segmentation.en_US
dc.description.degreeMaster of Science in Computer Science
dc.description.statementofresponsibilityFahim Ahmed Ifty
dc.format.extent60 pages
dc.identifier.otherID 23266007
dc.identifier.urihttp://hdl.handle.net/10361/27584
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectSemantic segmentationen_US
dc.subjectDeep learningen_US
dc.subjectResNet152en_US
dc.subjectData augmentationen_US
dc.subjectClass imbalanceen_US
dc.subjectCNNsen_US
dc.subjectOmniNeten_US
dc.subjectCamViden_US
dc.subjectTransformer refine blocken_US
dc.subjectConvolutional neural networks
dc.subjectUNeten_US
dc.subject.lcshDeep learning (Machine learning).
dc.subject.lcshNeural networks (Computer science).
dc.titleOmniNet: a hybrid deep learning framework for robust semantic segmentationen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
23266007_CSE.pdf
Size:
968.17 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: