Representation-aware unlearning via activation signatures: from suppression to knowledge-signature erasure

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorSadeque, Farig Yousuf
dc.contributor.authorBhuiyan, Md Rezaur Rahman
dc.contributor.authorMahmood, Syed Naveed
dc.contributor.authorKhondaker, Jareen Tasneem
dc.contributor.authorSakib, Md Sameer
dc.contributor.authorZaman, Tasfia
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-04-12T07:39:46Z
dc.date.available2026-04-12T07:39:46Z
dc.date.copyright2026
dc.date.issued2026-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 63-72).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.en_US
dc.description.abstractThe rapid development of large language models (LLMs) has outpaced regulations (e.g. GDPR) and ethical frameworks, raising concerns about privacy compliance, bias, misuse, misinformation, and legal adaptability. This makes the ability to selec- tively erase knowledge from LLMs critical. Despite significant developments, current unlearning methods are not able to segregate behavioral suppression and true knowl- edge removal, allowing latent capabilities to persist beneath surface-level refusals. In this paper, we address this challenge by introducing Knowledge Immunization Framework (KIF), a representation-aware architecture that differentiates between true erasure and obfuscation by operating on the internal activation signatures of the model, as opposed to surface-level outputs. KIF achieves near-oracle erasure (FQ ≈ 0.99 vs. 1.00) and utility preservation (MU = 0.62), effectively breaking the stability-erasure tradeoff that has constrained all prior work. Our observation shows that standard models exhibit scale-independent true erasure (<3% utility drift), while reasoning-prior models reveal fundamental architectural divergence. Our comprehensive dual-metric evaluation protocol, combining surface-level leakage with latent trace persistence, operationalizes the obfuscation - erasure distinction and enables the first systematic diagnosis of mechanism-level forgetting behavior across model families and scales.en_US
dc.description.degreeBachelor of Science in Computer Science and Engineering
dc.description.statementofresponsibilityMd Rezaur Rahman Bhuiyan
dc.description.statementofresponsibilitySyed Naveed Mahmood
dc.description.statementofresponsibilityJareen Tasneem Khondaker
dc.description.statementofresponsibilityMd Sameer Sakib
dc.description.statementofresponsibilityTasfia Zaman
dc.format.extent85 pages
dc.identifier.otherID 22301294
dc.identifier.otherID 22301257
dc.identifier.otherID 22301308
dc.identifier.otherID 24241344
dc.identifier.otherID 22301779
dc.identifier.urihttp://hdl.handle.net/10361/27862
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectLarge language modelen_US
dc.subjectKnowledge entanglementen_US
dc.subjectComputational overheaden_US
dc.subjectKnowledge immunization frameworken_US
dc.subjectKIFen_US
dc.subject.lcshMachine learning.
dc.subject.lcshKnowledge management.
dc.subject.lcshCopyright.
dc.titleRepresentation-aware unlearning via activation signatures: from suppression to knowledge-signature erasureen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
22301294, 22301257, 22301308, 24241344, 22301779_CSE.pdf
Size:
1.04 MB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: