Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Zero-shot detection of jailbreaking attempts in LLMs

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorShatabda, Swakkhar
dc.contributor.advisorChakrabarty, Amitabha
dc.contributor.authorRahman, Md. Hasib Ur
dc.contributor.authorAnkur, Arittra Paul
dc.contributor.authorZahin, Sadman
dc.contributor.authorFardin, Mominul Hoque
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-01-13T06:10:17Z
dc.date.available2026-01-13T06:10:17Z
dc.date.copyright2025
dc.date.issued2025-10
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 26-28).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.en_US
dc.description.abstractThe widespread deployment of Large Language Models (LLMs) has introduced significant safety challenges, notably the emergence of sophisticated ‘jailbreak’ attacks designed to bypass alignment measures and elicit harmful responses. While existing defenses often fail to generalize novel, zero-day attacks. we investigate the hypothesis that a classifier trained to a known distribution of attack patterns can achieve superior detection performance on entirely unseen adversarial prompts. We demonstrate that training on a specialized corpus of engineered safe prompts data that mirrors the structure and tonality of attacks—enhances the model’s ability to recognize conceptually similar yet novel threat vectors. When evaluated on a completely unseen challenge dataset of prompts confirmed to jailbreak state-of-theart models (including Grok-4, Grok-4 Heavy, and Gemini-2.5-Pro), our specialized detector improves accuracy from a baseline of 62.22% to 73.33%. These results, achieved with a compact training set, suggest that for rapidly evolving security threats like jailbreaking, targeted training with high-fidelity engineered data offers a more effective and resource-efficient defense mechanism than reliance on generalized, large-scale datasets.en_US
dc.description.degreeBachelor of Science in Computer Science and Engineering
dc.description.statementofresponsibilityMd. Hasib Ur Rahman
dc.description.statementofresponsibilityArittra Paul Ankur
dc.description.statementofresponsibilitySadman Zahin
dc.description.statementofresponsibilityMominul Hoque Fardin
dc.format.extent37 pages
dc.identifier.otherID 21201277
dc.identifier.otherID 21301101
dc.identifier.otherID 21101100
dc.identifier.otherID 20101518
dc.identifier.urihttp://hdl.handle.net/10361/27429
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectLarge language modelsen_US
dc.subjectJailbreak detectionen_US
dc.subjectAdversarial attacksen_US
dc.subjectEnsemble learningen_US
dc.subjectFeature engineeringen_US
dc.subjectSentence embeddingsen_US
dc.subject.lcshNatural language processing (Computer science).
dc.subject.lcshEnsemble learning (Machine learning).
dc.subject.lcshNatural language generation (Computer science)--Security measures.
dc.subject.lcshGenerative artificial intelligence--Security measures.
dc.subject.lcshComputer security--Risk assessment.
dc.titleZero-shot detection of jailbreaking attempts in LLMsen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
21201277, 21301101, 21101100, 20101518_CSE.pdf
Size:
472.46 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: