SPRINT: Sensitivity-guided pruning for inference-time adaptation of LLMs

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorHossain, Muhammad Iqbal
dc.contributor.authorDieyaz, Asief Iqbal
dc.contributor.authorHassan, Nahid
dc.contributor.authorYousuf, Faiyaz Bin
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-08-11T05:40:59Z
dc.date.available2026-08-11T05:40:59Z
dc.date.copyright2026
dc.date.issued2026-04
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 95-97).
dc.description.abstractLarge Language Models have transformed artificial intelligence, showing strong performance across a wide range of natural language processing tasks; however, their computational and memory demands make deployment on resource-constrained hardware, such as personal laptops, local servers, and edge devices, largely impractical. Existing compression techniques including pruning, quantization, knowledge distillation, and Low-Rank Adaptation reduce model overhead but apply a fixed compression level at deployment time, leaving them unable to respond to changes in prompt complexity, available memory, or latency requirements during inference. This work presents SPRINT, a dynamic hybrid inference framework that selects pruning intensity on a per-prompt basis at runtime. The system consists of three components: a Learned Complexity Router (LCR) built on a fine-tuned BERT-mini backbone that predicts each prompt’s sensitivity to pruning, a Double Deep Q-Network (DDQN) controller that selects pruning actions from a 17- option discrete space using a 10-dimensional state vector combining hardware telemetry, router scores, and early backbone signals, and a structural pruning engine that physically removes transformer layers to produce real latency reductions. Experiments were conducted on Llama-2-7B using a 10,000-prompt dataset drawn equally from GSM8K, MBPP, WikiText-2, MMLU, and BoolQ. The LCR achieved a Spearman rank correlation of ρ = 0.797 (95% CI: [0.779, 0.817]) and R2 = 0.633 against oracle sensitivity labels on the held-out test set, exceeding the target threshold of ρ ≥ 0.70 across all five benchmark domains. Across 2,000 held-out test episodes, the DDQN controller reduced average inference time from 1,287.61 ms to 875.39 ms, yielding a 32.0% average speedup. Parameter count dropped from 4,714.3 MB to 3,113.5 MB, a reduction of 1,600.8 MB (34.0%), while peak VRAM consumption remained stable at approximately 4.75 GB. Total routing and action-selection overhead averaged just 17.77 ms, corresponding to approximately 2.0% of total latency. Comparison against SparseGPT, Wanda, and LLM Pruner confirmed that unstructured weight sparsity does not reliably convert to latency reduction on standard GPU hardware, whereas SPRINT’s structural layer removal yields predictable speedups without requiring sparse kernel support.
dc.description.degreeBachelor of Science in Computer Science and Engineering
dc.description.statementofresponsibilityAsief Iqbal Dieyaz
dc.description.statementofresponsibilityNahid Hassan
dc.description.statementofresponsibilityFaiyaz Bin Yousuf
dc.format.extent106 pages
dc.identifier.otherID 23341081
dc.identifier.otherID 22101822
dc.identifier.otherID 22101845
dc.identifier.urihttps://hdl.handle.net/10361/28913
dc.language.isoen_US
dc.publisherBRAC University
dc.rightsAttribution-NonCommercial-NoDerivatives 4.0 Internationalen
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/4.0/
dc.subjectLarge language models
dc.subjectLLMs
dc.subjectStructural pruning
dc.subjectMachine learning
dc.subjectResource-aware systems
dc.subjectDDQN
dc.subjectKnowledge distillation
dc.subjectAdaptive inference
dc.subjectEdge deployment
dc.subjectDynamic optimization
dc.subjectNatural language processing
dc.subject.lcshNatural language processing (Computer science).
dc.subject.lcshArtificial intelligence.
dc.subject.lcshReinforcement learning.
dc.titleSPRINT: Sensitivity-guided pruning for inference-time adaptation of LLMs
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
22101822, 23341081, 22101845_CSE.pdf
Size:
1.8 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: