Dynamic gated distillation: dual-teacher framework for jailbreak prevention in LLMs with preserved reasoning

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorChakrabarty, Amitabha
dc.contributor.authorSarker, Rwittik
dc.contributor.authorRafat, Nabil Walid
dc.contributor.authorOmi, Anika Tabassum
dc.contributor.authorAuvi, Ferdous Ahmed
dc.contributor.authorRahman, Zabid
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-04-12T10:44:40Z
dc.date.available2026-04-12T10:44:40Z
dc.date.copyright2026
dc.date.issued2026-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 41-45).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.en_US
dc.description.abstractKnowledge Distillation (KD) can be used to reduce large language models into smaller and more efficient models, although traditional knowledge distillation can preserve unsafe guidance on behav- ioral instructions, so small language models (SLMs) are vulnerable to jailbreak attacks. Moreover, using traditional knowledge distillation to finetune in terms of jailbreak robustness often degrades general reasoning performance. This creates a tough challenge to preserve jailbreak robustness as well as general reasoning at the same time. So often traditional knowledge distillation falls short to tackle two downstream tasks in one SFT. We introduce Dynamic Gated Distillation (DGD), a novel framework to which harm-score prediction has been added to enhance safety-utility alignment in the process of a dual teacher distillation. Comparative analysis shows that our DGD framework manages to retain almost all of its general alignment (MMLU score 42.13%) after distillation, com- pared with the base student model (MMLU score 42.81%), while reducing the attack success rate (ASR) below 5%, demonstrating the framework’s effectiveness in achieving both safety and utility in small language models (SLMs).en_US
dc.description.degreeBachelor of Science in Computer Science and Engineering
dc.description.statementofresponsibilityRwittik Sarker
dc.description.statementofresponsibilityNabil Walid Rafat
dc.description.statementofresponsibilityAnika Tabassum Omi
dc.description.statementofresponsibilityFerdous Ahmed Auvi
dc.description.statementofresponsibilityZabid Rahman
dc.format.extent55 pages
dc.identifier.otherID 22101634
dc.identifier.otherID 22101613
dc.identifier.otherID 22101794
dc.identifier.otherID 22101583
dc.identifier.otherID 24141085
dc.identifier.urihttp://hdl.handle.net/10361/27869
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectKnowledge distillationen_US
dc.subjectGated weighten_US
dc.subjectJailbreak promptsen_US
dc.subjectLarge language modelen_US
dc.subject.lcshDistillation.
dc.subject.lcshKnowledge.
dc.subject.lcshMachine learning.
dc.titleDynamic gated distillation: dual-teacher framework for jailbreak prevention in LLMs with preserved reasoningen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
22101634, 22101613, 22101794, 22101583, 24141085_CSE.pdf
Size:
494.16 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: