Punolikhon: a deep learning approach for handwritten text recognition of Bengali and English documents using a lightweight model
Loading...
Date
Publisher
BRAC University
Citation
Abstract
It is very difficult for computers to read and extract data from the handwritten form of
documents in bilingual models like Bengali and English. The condition becomes much
more complex for Bengali scripts as it has numerous complex scripts, integrated letters,
numerous styles, which makes the recognition process way more complex. This research
focuses on a deep learning oriented model named Punolikon, a lightweight bilingual Optical
Character Recognition Model which helps to identify both the Bengali and English
handwritten form of documents efficiently and effectively. The model applies a novel hybrid
deep learning model which includes Spatial Transformer Networks (STN) for input
normalization, ResNet-style bottleneck blocks with skip connections for channel reduction,
MobileNetV3-Small backbone for effective feature extraction and Transformer encoders
with multi-head self-attention for robust model sequence. The recognition module
handles conjunct consonants, vowel signs and Unicode normalization problems by using a
grapheme-based tokenization technique created especially for the complex Bengali letter.
For ensuring clean and clear text portion this model also relies on CLAHE along with
graysclale at preprocessing period. This hybrid model achieves 84.3% word accuracy and
9.87% Character Error Rate on overall text recognition using a test set of 18,953 word
samples. Fixing Unicode Normalization improved accuracy by 11%. The model shows
strong performance with 88.4% character accuracy and 92.08% average prediction confidence.
The detection part uses YOLOv8n for word localization and for the recognition
pipeline applies a hybrid STN-CNN-Transformer structure developed with CTC loss to
handle word crops. Finally, this integrated architecture helps the model to gain satisfactory
rate of accuracy in case of recognizing both Bengali (85.3%) and English (83.3%)
handwriting by keeping the system robust and effective. Also this provides a comprehensive
solution for digitising handwritten and printed historical Bengali-English records,
offers administration digitisation and academical resources preservations.
Description
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 53-58).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.
Includes bibliographical references (pages 53-58).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.
Publisher Link
Type
Thesis