Alam, Md. Golam RabiulIfty, Fahim Ahmed2026-03-042026-03-0420252025-10ID 23266007http://hdl.handle.net/10361/27584Cataloged from PDF version of thesis.Includes bibliographical references (pages 45-47).This thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science, 2025.Semantic segmentation empowers computers to interpret visual scenes in a structured and meaningful manner, accurately delineating the precise boundaries of every object in an image at the pixel level. However, existing CNN-based models struggle with long-range context, transformer-based approaches are computationally heavy and often miss local detail, and many hybrid designs neglect explicit multi-scale features or refined attention, while also suffering from class imbalance and noisy annotations. This research proposes an encoder–decoder architecture that combines a ResNet152 backbone, Atrous Spatial Pyramid Pooling (ASPP), and a Transformer Refine Block, leveraging the strengths of convolutional features and lightweight self-attention to capture both multi-scale semantics and long-range dependencies. In the decoder, a UNet-style upsampling path with skip connections and CBAM attention preserves spatial detail and refines boundaries, yielding coherent pixel-level predictions across scales. The training pipeline applies label cleaning and robust augmentation (MixUp, CutMix) and optimizes a composite loss (Lovasz–Softmax, Dice Loss, Boundary Loss) to counter label noise and the long-tail distribution. Evaluated on the CamVid benchmark, our method achieves 86.53% mean IoU and 95.99% pixel accuracy, outperforming recent efficient residual attention networks while prioritizing segmentation quality over inference speed. These results demonstrate that combining multi-scale convolutional features with lightweight self-attention and boundary-aware optimization delivers state-of-the-art accuracy under noisy labels and severe class imbalance in urban scene segmentation.60 pagesenBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.Semantic segmentationDeep learningResNet152Data augmentationClass imbalanceCNNsOmniNetCamVidTransformer refine blockConvolutional neural networksUNetDeep learning (Machine learning).Neural networks (Computer science).OmniNet: a hybrid deep learning framework for robust semantic segmentationThesis