Alam, Md. Golam RabiulShatabda, SwakkharIslam, ShakerSami, Md. Ashraful IslamRohan, Amin MohammadDeb, Jhishan2026-08-092026-08-0920262026-01ID 22341007ID 22301118ID 22301338ID 22301357https://hdl.handle.net/10361/28840This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.Cataloged from PDF version of thesis.Includes bibliographical references (pages 67-69).Long read sequencing is a DNA sequencing technique that makes it possible to sequence long DNA fragments, offering an unprecedented opportunity to resolve structural variants (SVs) in genomes. Structural variants (SVs) are changes in DNA that play an important role in genome diversity, evolution, and disease. Detecting SVs in the human genome is challenging because of differences in genome structure, complexity, and limited labeled data. Existing long read structural variant detection methods depend on predefined rules or heuristic strategies, which do not fully capture the intricate nature of SV signatures. To overcome these limitations, we introduce a transformer driven model for analyzing structural variants in the human genome called SVBERT. Our approach first extracts SV signatures from sequence alignments and assembles local regions to generate paired sequence inputs. Local alignment features are processed by a modified convolutional neural network (CNN) encoder, while BERT generates context aware sequence embeddings. A fusion module then combines these features using cross attention followed by a transformer encoder. Finally, specialized prediction heads perform classification, breakpoint regression, genotype calling, and confidence scoring. Post processing with confidence based filtering produces high quality structural variant calls. Validation across human genome sequencing datasets shows improved detection of various SV types. These results demonstrate the strong potential of transformer based genomic language models for advancing SV analysis in both research and practical applications.70 pagesen-USAttribution-NonCommercial-NoDerivatives 4.0 InternationalBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.http://creativecommons.org/licenses/by-nc-nd/4.0/Long read sequencingGenomic language modelDNA sequencingVariant detectionStructural variantsHuman genomeComputational biology.Machine learning.Nucleotide sequence.Neural networks (Computer science).Bioinformatics.Structural variant analysis in human genome using transformer based genomic language modelThesis