Gated ResNet architecture for transformer multi-head attention block

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorSadeque, Farig Yousuf
dc.contributor.advisorAbedin, Jawaril Munshad
dc.contributor.authorHassan, Ahnaf
dc.contributor.authorMahde, Mubtasim Fuad
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2025-06-16T04:14:06Z
dc.date.available2025-06-16T04:14:06Z
dc.date.copyright2025
dc.date.issued2025-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 56-60).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2025.en_US
dc.description.abstractTransformer model architectures show excellent performance over machine learning- related tasks, primarily focusing on natural language processing by removing the need for recurrence and convolutional techniques. Residual networks (ResNet) are a core component for transformer models to retain long-term dependencies for sequence-to-sequence-related tasks. This paper introduces a new ResNet for the at- tention layer of a transformer model. Through a ResNet connection between layers of multi-head blocks, we flow information from one layer to the next while keeping the number of parameters the same. We also explore two-layer deep connections that retain long-term dependencies even more than one layer. Furthermore, we im- plement a gating mechanism on the ResNet that will selectively allow less redundant information to flow through and ensure that the gradient convergence of the model can be accelerated. In this study, we intend to prove that the performance of a model can increase even if the parameter increases due to introducing the gated unit which is insignificant in comparison to the overall number of parameters. Therefore, en- hancing the model’s performance on seq2seq tasks without demanding additional training time and memory by adjusting a few trainable parameters introduced for the gating mechanism is the primary goal of the paper. Our model showed promis- ing results, especially on long-term dependent seq2seq tasks, by achieving a better performance score as well as maintaining similar efficiency.en_US
dc.description.degreeBachelor of Science in Computer Science
dc.description.statementofresponsibilityAhnaf Hassan
dc.description.statementofresponsibilityMubtasim Fuad Mahde
dc.format.extent61 pages
dc.identifier.otherID 21341009
dc.identifier.otherID 24341116
dc.identifier.urihttp://hdl.handle.net/10361/26030
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectTransformer modelen_US
dc.subjectResidual networks (ResNet)en_US
dc.subjectAttention layeren_US
dc.subjectMulti- head blocksen_US
dc.subjectLong-term dependenciesen_US
dc.subjectSequence-to-sequence tasksen_US
dc.subjectGating mechanismen_US
dc.subjectGradient convergenceen_US
dc.subjectParameter increaseen_US
dc.subjectMemory reductionen_US
dc.subjectSeq2seq tasksen_US
dc.subjectTraining stabilityen_US
dc.subjectSparse attention scoresen_US
dc.subjectModel performanceen_US
dc.subjectInformation flowen_US
dc.subjectBLEU scoreen_US
dc.subject.lcshComputational intelligence
dc.titleGated ResNet architecture for transformer multi-head attention blocken_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
24341116_21341009_CSE.pdf
Size:
3.86 MB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: