Anchor-guided repair: a defense mechanism for enhancing stability of compromised pretrained language models against low-precision and weight noise attacks
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Hossain, Muhammad Iqbal | |
| dc.contributor.author | Rohan, Abrar Mahir | |
| dc.contributor.author | Khan, Nafiz | |
| dc.contributor.author | Tanni, Tahsin Tajwar | |
| dc.contributor.author | Fardin, Fuad | |
| dc.contributor.author | Bushra, Anika | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-04-19T09:16:52Z | |
| dc.date.available | 2026-04-19T09:16:52Z | |
| dc.date.copyright | 2026 | |
| dc.date.issued | 2026-01 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 61-62). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026. | en_US |
| dc.description.abstract | Large Language Models (LLMs) are increasingly released as open-source and are susceptible to post-release attacks, including weight noise injection and low-precision quantization. Such attacks often have a significant negative effect on model stability and performance, augmenting perplexity and creating unreliable behavior. In practice, users might not have access to original clean weights and may unknowingly download compromised models. This thesis presents a defense mechanism called Anchor-Guided Repair, which aims to improve the stability of weakened LLMs against weight noise and low-precision attacks. The proposed method optimizes a fine-tuned attacked model on clean textual data while mitigating parameter changes via an anchor loss, which penalizes the difference from a clean, task-adapted baseline. This approach reduces instability by optimizing a composite task involving language modeling loss and anchor regularization, without distorting the knowledge acquired by the original model. The proposed method was put to the test across a range of model architectures and attack setups, including low-precision and Gaussian weight noise. The experimental results clearly show that Anchor-Guided Repair outperforms the attacked models consistently, keeping a favorable state of desirable conditions. While the defended model performs slightly below the clean baseline, it is significantly better than the compromised model. This emphasizes the power of anchoring in stabilizing large models without accessing proprietary training data. Altogether, Anchor-Guided Repair is a viable post-deployment defense for weightlevel attacks, promoting safer and more trustworthy model reuse in open-source settings. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Abrar Mahir Rohan | |
| dc.description.statementofresponsibility | Nafiz Khan | |
| dc.description.statementofresponsibility | Tahsin Tajwar Tanni | |
| dc.description.statementofresponsibility | Fuad Fardin | |
| dc.description.statementofresponsibility | Anika Bushra | |
| dc.format.extent | 62 pages | |
| dc.identifier.other | ID 22101267 | |
| dc.identifier.other | ID 22101045 | |
| dc.identifier.other | ID 22101744 | |
| dc.identifier.other | ID 22101027 | |
| dc.identifier.other | ID 21201068 | |
| dc.identifier.uri | http://hdl.handle.net/10361/27945 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Quantization attacks | en_US |
| dc.subject | Large language models | en_US |
| dc.subject | Weight noise injection | en_US |
| dc.subject | Backdoor detection | en_US |
| dc.subject | Llow-precision quantization | en_US |
| dc.subject | Anchor-guided repair | en_US |
| dc.subject.lcsh | Natural language generation (Computer science)--Security measures. | |
| dc.subject.lcsh | Generative artificial intelligence--Security measures. | |
| dc.subject.lcsh | Deep learning (Machine learning)--Security measures. | |
| dc.subject.lcsh | Machine learning--Mathematical models. | |
| dc.subject.lcsh | Neural networks (Computer science). | |
| dc.subject.lcsh | Computer software--Reliability. | |
| dc.title | Anchor-guided repair: a defense mechanism for enhancing stability of compromised pretrained language models against low-precision and weight noise attacks | en_US |
| dc.type | Thesis | en_US |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- 22101267, 22101045, 22101744, 22101027, 21201068_CSE.pdf
- Size:
- 922.82 KB
- Format:
- Adobe Portable Document Format
- Description:
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: