AI-driven context-aware trimming and segmentation of educational video content for focused learning

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorRabiul Alam, Md. Golam
dc.contributor.advisorReza, Md Tanzim
dc.contributor.authorNabil, Jonayed Kader
dc.contributor.authorChowdhury, Tasnim Aziz
dc.contributor.authorHarun, Sayed Ilham Azhar
dc.contributor.authorMozumder, Rafi Ehtesham
dc.contributor.authorTasnim, Nuzhat
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-04-26T06:01:32Z
dc.date.available2026-04-26T06:01:32Z
dc.date.copyright2026
dc.date.issued2026-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 47-50).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.en_US
dc.description.abstractRecorded class lectures and tutorials are essential learning resources, but unedited raw videos generally contain significant non-instructional content, such as silence, administrative discussions, off-topic instructions, making effective navigation and revision difficult. This study presents a framework which is open-source, contextaware and can automatically detect and remove irrelevant or off-topic segments from long, unedited educational videos. Most of the existing related solutions rely heavily on large, closed-source models or fixed-length video segmentation with no context awareness, which are computationally expensive and prone to breaking semantic continuity. To address these limitations, we propose a modular, four-stage orchestration pipeline designed to run entirely on consumer-grade GPUs using small, open-weight models. The framework uses automatic speech recognition (Parakeet TDT 0.6B) to perform dynamic, sentence-level video segmentation, ensuring that each segment represents a complete semantic unit. A lightweight vision language model then generates structured textual descriptions for each segment using both visual frames and subtitles, augmented with rolling local context to preserve chronological coherence. Finally, a small language model (Qwen3-4B-instruct) is used to perform classification on the generated descriptions rather than on raw video input. The proposed orchestrator pipeline achieves 96.85% accuracy, 89.04% precision, 92.77% recall, and 90.86% F1 for irrelevant-segment detection on a human-annotated benchmark of 20 educational videos (total duration 09:13:47). This results in a significant improvement in content preservation over a one-shot Gemini-3 baseline, with precision increasing from 45.98% to 89.04% and accuracy from 81.63% to 96.85%. The entire system runs locally via quantized inference, offering a practical and privacypreserving alternative to cloud-based solutions.en_US
dc.description.degreeBachelor of Science in Computer Science and Engineering
dc.description.statementofresponsibilityJonayed Kader Nabil
dc.description.statementofresponsibilityTasnim Aziz Chowdhury
dc.description.statementofresponsibilitySayed Ilham Azhar Harun
dc.description.statementofresponsibilityRafi Ehtesham Mozumder
dc.description.statementofresponsibilityNuzhat Tasnim
dc.format.extent53 pages
dc.identifier.otherID 22101100
dc.identifier.otherID 22101164
dc.identifier.otherID 22101262
dc.identifier.otherID 22301489
dc.identifier.otherID 22101167
dc.identifier.urihttp://hdl.handle.net/10361/28064
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectEducational video analysisen_US
dc.subjectIrrelevant content detectionen_US
dc.subjectVision-language modelsen_US
dc.subjectSpeech recognitionen_US
dc.subjectVideo processingen_US
dc.subjectMachine learningen_US
dc.subject.lcshEducational technology.
dc.subject.lcshComputer-assisted instruction.
dc.subject.lcshEducational technology.
dc.subject.lcshSpeech processing systems.
dc.titleAI-driven context-aware trimming and segmentation of educational video content for focused learningen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
22101100, 22101164, 22101262, 22301489, 22101167_CSE.pdf
Size:
874.36 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: