Zero-shot detection of jailbreaking attempts in LLMs
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Shatabda, Swakkhar | |
| dc.contributor.advisor | Chakrabarty, Amitabha | |
| dc.contributor.author | Rahman, Md. Hasib Ur | |
| dc.contributor.author | Ankur, Arittra Paul | |
| dc.contributor.author | Zahin, Sadman | |
| dc.contributor.author | Fardin, Mominul Hoque | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-01-13T06:10:17Z | |
| dc.date.available | 2026-01-13T06:10:17Z | |
| dc.date.copyright | 2025 | |
| dc.date.issued | 2025-10 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 26-28). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025. | en_US |
| dc.description.abstract | The widespread deployment of Large Language Models (LLMs) has introduced significant safety challenges, notably the emergence of sophisticated ‘jailbreak’ attacks designed to bypass alignment measures and elicit harmful responses. While existing defenses often fail to generalize novel, zero-day attacks. we investigate the hypothesis that a classifier trained to a known distribution of attack patterns can achieve superior detection performance on entirely unseen adversarial prompts. We demonstrate that training on a specialized corpus of engineered safe prompts data that mirrors the structure and tonality of attacks—enhances the model’s ability to recognize conceptually similar yet novel threat vectors. When evaluated on a completely unseen challenge dataset of prompts confirmed to jailbreak state-of-theart models (including Grok-4, Grok-4 Heavy, and Gemini-2.5-Pro), our specialized detector improves accuracy from a baseline of 62.22% to 73.33%. These results, achieved with a compact training set, suggest that for rapidly evolving security threats like jailbreaking, targeted training with high-fidelity engineered data offers a more effective and resource-efficient defense mechanism than reliance on generalized, large-scale datasets. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Md. Hasib Ur Rahman | |
| dc.description.statementofresponsibility | Arittra Paul Ankur | |
| dc.description.statementofresponsibility | Sadman Zahin | |
| dc.description.statementofresponsibility | Mominul Hoque Fardin | |
| dc.format.extent | 37 pages | |
| dc.identifier.other | ID 21201277 | |
| dc.identifier.other | ID 21301101 | |
| dc.identifier.other | ID 21101100 | |
| dc.identifier.other | ID 20101518 | |
| dc.identifier.uri | http://hdl.handle.net/10361/27429 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Large language models | en_US |
| dc.subject | Jailbreak detection | en_US |
| dc.subject | Adversarial attacks | en_US |
| dc.subject | Ensemble learning | en_US |
| dc.subject | Feature engineering | en_US |
| dc.subject | Sentence embeddings | en_US |
| dc.subject.lcsh | Natural language processing (Computer science). | |
| dc.subject.lcsh | Ensemble learning (Machine learning). | |
| dc.subject.lcsh | Natural language generation (Computer science)--Security measures. | |
| dc.subject.lcsh | Generative artificial intelligence--Security measures. | |
| dc.subject.lcsh | Computer security--Risk assessment. | |
| dc.title | Zero-shot detection of jailbreaking attempts in LLMs | en_US |
| dc.type | Thesis | en_US |