Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Anchor-guided repair: a defense mechanism for enhancing stability of compromised pretrained language models against low-precision and weight noise attacks

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorHossain, Muhammad Iqbal
dc.contributor.authorRohan, Abrar Mahir
dc.contributor.authorKhan, Nafiz
dc.contributor.authorTanni, Tahsin Tajwar
dc.contributor.authorFardin, Fuad
dc.contributor.authorBushra, Anika
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-04-19T09:16:52Z
dc.date.available2026-04-19T09:16:52Z
dc.date.copyright2026
dc.date.issued2026-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 61-62).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.en_US
dc.description.abstractLarge Language Models (LLMs) are increasingly released as open-source and are susceptible to post-release attacks, including weight noise injection and low-precision quantization. Such attacks often have a significant negative effect on model stability and performance, augmenting perplexity and creating unreliable behavior. In practice, users might not have access to original clean weights and may unknowingly download compromised models. This thesis presents a defense mechanism called Anchor-Guided Repair, which aims to improve the stability of weakened LLMs against weight noise and low-precision attacks. The proposed method optimizes a fine-tuned attacked model on clean textual data while mitigating parameter changes via an anchor loss, which penalizes the difference from a clean, task-adapted baseline. This approach reduces instability by optimizing a composite task involving language modeling loss and anchor regularization, without distorting the knowledge acquired by the original model. The proposed method was put to the test across a range of model architectures and attack setups, including low-precision and Gaussian weight noise. The experimental results clearly show that Anchor-Guided Repair outperforms the attacked models consistently, keeping a favorable state of desirable conditions. While the defended model performs slightly below the clean baseline, it is significantly better than the compromised model. This emphasizes the power of anchoring in stabilizing large models without accessing proprietary training data. Altogether, Anchor-Guided Repair is a viable post-deployment defense for weightlevel attacks, promoting safer and more trustworthy model reuse in open-source settings.en_US
dc.description.degreeBachelor of Science in Computer Science and Engineering
dc.description.statementofresponsibilityAbrar Mahir Rohan
dc.description.statementofresponsibilityNafiz Khan
dc.description.statementofresponsibilityTahsin Tajwar Tanni
dc.description.statementofresponsibilityFuad Fardin
dc.description.statementofresponsibilityAnika Bushra
dc.format.extent62 pages
dc.identifier.otherID 22101267
dc.identifier.otherID 22101045
dc.identifier.otherID 22101744
dc.identifier.otherID 22101027
dc.identifier.otherID 21201068
dc.identifier.urihttp://hdl.handle.net/10361/27945
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectQuantization attacksen_US
dc.subjectLarge language modelsen_US
dc.subjectWeight noise injectionen_US
dc.subjectBackdoor detectionen_US
dc.subjectLlow-precision quantizationen_US
dc.subjectAnchor-guided repairen_US
dc.subject.lcshNatural language generation (Computer science)--Security measures.
dc.subject.lcshGenerative artificial intelligence--Security measures.
dc.subject.lcshDeep learning (Machine learning)--Security measures.
dc.subject.lcshMachine learning--Mathematical models.
dc.subject.lcshNeural networks (Computer science).
dc.subject.lcshComputer software--Reliability.
dc.titleAnchor-guided repair: a defense mechanism for enhancing stability of compromised pretrained language models against low-precision and weight noise attacksen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
22101267, 22101045, 22101744, 22101027, 21201068_CSE.pdf
Size:
922.82 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: