AI-driven context-aware trimming and segmentation of educational video content for focused learning
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Rabiul Alam, Md. Golam | |
| dc.contributor.advisor | Reza, Md Tanzim | |
| dc.contributor.author | Nabil, Jonayed Kader | |
| dc.contributor.author | Chowdhury, Tasnim Aziz | |
| dc.contributor.author | Harun, Sayed Ilham Azhar | |
| dc.contributor.author | Mozumder, Rafi Ehtesham | |
| dc.contributor.author | Tasnim, Nuzhat | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-04-26T06:01:32Z | |
| dc.date.available | 2026-04-26T06:01:32Z | |
| dc.date.copyright | 2026 | |
| dc.date.issued | 2026-01 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 47-50). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026. | en_US |
| dc.description.abstract | Recorded class lectures and tutorials are essential learning resources, but unedited raw videos generally contain significant non-instructional content, such as silence, administrative discussions, off-topic instructions, making effective navigation and revision difficult. This study presents a framework which is open-source, contextaware and can automatically detect and remove irrelevant or off-topic segments from long, unedited educational videos. Most of the existing related solutions rely heavily on large, closed-source models or fixed-length video segmentation with no context awareness, which are computationally expensive and prone to breaking semantic continuity. To address these limitations, we propose a modular, four-stage orchestration pipeline designed to run entirely on consumer-grade GPUs using small, open-weight models. The framework uses automatic speech recognition (Parakeet TDT 0.6B) to perform dynamic, sentence-level video segmentation, ensuring that each segment represents a complete semantic unit. A lightweight vision language model then generates structured textual descriptions for each segment using both visual frames and subtitles, augmented with rolling local context to preserve chronological coherence. Finally, a small language model (Qwen3-4B-instruct) is used to perform classification on the generated descriptions rather than on raw video input. The proposed orchestrator pipeline achieves 96.85% accuracy, 89.04% precision, 92.77% recall, and 90.86% F1 for irrelevant-segment detection on a human-annotated benchmark of 20 educational videos (total duration 09:13:47). This results in a significant improvement in content preservation over a one-shot Gemini-3 baseline, with precision increasing from 45.98% to 89.04% and accuracy from 81.63% to 96.85%. The entire system runs locally via quantized inference, offering a practical and privacypreserving alternative to cloud-based solutions. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Jonayed Kader Nabil | |
| dc.description.statementofresponsibility | Tasnim Aziz Chowdhury | |
| dc.description.statementofresponsibility | Sayed Ilham Azhar Harun | |
| dc.description.statementofresponsibility | Rafi Ehtesham Mozumder | |
| dc.description.statementofresponsibility | Nuzhat Tasnim | |
| dc.format.extent | 53 pages | |
| dc.identifier.other | ID 22101100 | |
| dc.identifier.other | ID 22101164 | |
| dc.identifier.other | ID 22101262 | |
| dc.identifier.other | ID 22301489 | |
| dc.identifier.other | ID 22101167 | |
| dc.identifier.uri | http://hdl.handle.net/10361/28064 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Educational video analysis | en_US |
| dc.subject | Irrelevant content detection | en_US |
| dc.subject | Vision-language models | en_US |
| dc.subject | Speech recognition | en_US |
| dc.subject | Video processing | en_US |
| dc.subject | Machine learning | en_US |
| dc.subject.lcsh | Educational technology. | |
| dc.subject.lcsh | Computer-assisted instruction. | |
| dc.subject.lcsh | Educational technology. | |
| dc.subject.lcsh | Speech processing systems. | |
| dc.title | AI-driven context-aware trimming and segmentation of educational video content for focused learning | en_US |
| dc.type | Thesis | en_US |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- 22101100, 22101164, 22101262, 22301489, 22101167_CSE.pdf
- Size:
- 874.36 KB
- Format:
- Adobe Portable Document Format
- Description:
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: