Real time closed captioning for Bengali multimedia
Loading...
Date
Publisher
BRAC University
Citation
Abstract
The widespread use of Bengali multimedia content on different platforms increases
the need for an accessibility tool to cater to the Bengali content audience. Despite
having tools for closed captioning pre recorded videos, there is still a need for a
real time captioner to address the problem of captioning live programs. Closed
captioning of live programs can make these contents more accessible to a wide
range of people, including hard of hearing people and those who are not accustomed
to the Bengali language. Due to environmental noise interference, temporal
synchronization requirements, and overlap artifacts in streaming transcription systems,
real-time closed captioning for Bengali multimedia content faces critical challenges.
This paper presents a comprehensive framework for robust Bengali closed
captioning that integrates a novel architecture optimized for multimedia applications.
Our approach combines VAD-enhanced audio preprocessing, Silero-based filtering
with Bengali-optimized thresholds (speech probability 0.25, ratio 0.15), and
adaptive spectral noise suppression using conservative subtraction (=1.5) with 30%
spectral floor protection to preserve conjunct consonants and tonal variations critical
for Bengali intelligibility. To establish optimal acoustic model foundation, we systematically
evaluated custom -xls-r-300m and whiper-small fine-tuning on 20 hours
of Mozilla CommonVoice data versus pre-trained alternatives, demonstrating that
larger training datasets (300+ hours) provide 35-45% superior performance for multimedia
captioning scenarios. Our overlap prevention framework employs sliding
window deduplication and cross-stride output comparison with Bengali morphological
fuzzy matching. While it filters many overlap, there is room for improvement
in overlapping prevention. The system maintains sub-second latency with minimal
computational overhead, advancing accessible Bengali multimedia consumption for
hearing-impaired communities and multilingual audiences. However, further work
is needed on handling real-life multimedia challenges including speaker variations,
tonal variation due to frequent emotional changes, background music interference,
and cross-talk scenarios to achieve broadcast-quality transcription accuracy for diverse
multimedia content.
Description
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 43-44).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.
Includes bibliographical references (pages 43-44).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.
Publisher Link
Type
Thesis