Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Real time closed captioning for Bengali multimedia

Citation

Abstract

The widespread use of Bengali multimedia content on different platforms increases the need for an accessibility tool to cater to the Bengali content audience. Despite having tools for closed captioning pre recorded videos, there is still a need for a real time captioner to address the problem of captioning live programs. Closed captioning of live programs can make these contents more accessible to a wide range of people, including hard of hearing people and those who are not accustomed to the Bengali language. Due to environmental noise interference, temporal synchronization requirements, and overlap artifacts in streaming transcription systems, real-time closed captioning for Bengali multimedia content faces critical challenges. This paper presents a comprehensive framework for robust Bengali closed captioning that integrates a novel architecture optimized for multimedia applications. Our approach combines VAD-enhanced audio preprocessing, Silero-based filtering with Bengali-optimized thresholds (speech probability 0.25, ratio 0.15), and adaptive spectral noise suppression using conservative subtraction (=1.5) with 30% spectral floor protection to preserve conjunct consonants and tonal variations critical for Bengali intelligibility. To establish optimal acoustic model foundation, we systematically evaluated custom -xls-r-300m and whiper-small fine-tuning on 20 hours of Mozilla CommonVoice data versus pre-trained alternatives, demonstrating that larger training datasets (300+ hours) provide 35-45% superior performance for multimedia captioning scenarios. Our overlap prevention framework employs sliding window deduplication and cross-stride output comparison with Bengali morphological fuzzy matching. While it filters many overlap, there is room for improvement in overlapping prevention. The system maintains sub-second latency with minimal computational overhead, advancing accessible Bengali multimedia consumption for hearing-impaired communities and multilingual audiences. However, further work is needed on handling real-life multimedia challenges including speaker variations, tonal variation due to frequent emotional changes, background music interference, and cross-talk scenarios to achieve broadcast-quality transcription accuracy for diverse multimedia content.

Description

Cataloged from PDF version of thesis.
Includes bibliographical references (pages 43-44).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.

Publisher Link

Type

Thesis