Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Punolikhon: a deep learning approach for handwritten text recognition of Bengali and English documents using a lightweight model

Citation

Abstract

It is very difficult for computers to read and extract data from the handwritten form of documents in bilingual models like Bengali and English. The condition becomes much more complex for Bengali scripts as it has numerous complex scripts, integrated letters, numerous styles, which makes the recognition process way more complex. This research focuses on a deep learning oriented model named Punolikon, a lightweight bilingual Optical Character Recognition Model which helps to identify both the Bengali and English handwritten form of documents efficiently and effectively. The model applies a novel hybrid deep learning model which includes Spatial Transformer Networks (STN) for input normalization, ResNet-style bottleneck blocks with skip connections for channel reduction, MobileNetV3-Small backbone for effective feature extraction and Transformer encoders with multi-head self-attention for robust model sequence. The recognition module handles conjunct consonants, vowel signs and Unicode normalization problems by using a grapheme-based tokenization technique created especially for the complex Bengali letter. For ensuring clean and clear text portion this model also relies on CLAHE along with graysclale at preprocessing period. This hybrid model achieves 84.3% word accuracy and 9.87% Character Error Rate on overall text recognition using a test set of 18,953 word samples. Fixing Unicode Normalization improved accuracy by 11%. The model shows strong performance with 88.4% character accuracy and 92.08% average prediction confidence. The detection part uses YOLOv8n for word localization and for the recognition pipeline applies a hybrid STN-CNN-Transformer structure developed with CTC loss to handle word crops. Finally, this integrated architecture helps the model to gain satisfactory rate of accuracy in case of recognizing both Bengali (85.3%) and English (83.3%) handwriting by keeping the system robust and effective. Also this provides a comprehensive solution for digitising handwritten and printed historical Bengali-English records, offers administration digitisation and academical resources preservations.

Description

Cataloged from PDF version of thesis.
Includes bibliographical references (pages 53-58).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.

Publisher Link

Type

Thesis