Punolikhon: a deep learning approach for handwritten text recognition of Bengali and English documents using a lightweight model
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Mukta, Jannatun Noor | |
| dc.contributor.author | Zaman, Sanjana | |
| dc.contributor.author | Noman, Md. Kawser Alam | |
| dc.contributor.author | Raiyan, Khandoker Sakib | |
| dc.contributor.author | Hossain, Md. Sulaiman | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-04-19T06:04:31Z | |
| dc.date.available | 2026-04-19T06:04:31Z | |
| dc.date.copyright | 2026 | |
| dc.date.issued | 2026-02 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 53-58). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026. | en_US |
| dc.description.abstract | It is very difficult for computers to read and extract data from the handwritten form of documents in bilingual models like Bengali and English. The condition becomes much more complex for Bengali scripts as it has numerous complex scripts, integrated letters, numerous styles, which makes the recognition process way more complex. This research focuses on a deep learning oriented model named Punolikon, a lightweight bilingual Optical Character Recognition Model which helps to identify both the Bengali and English handwritten form of documents efficiently and effectively. The model applies a novel hybrid deep learning model which includes Spatial Transformer Networks (STN) for input normalization, ResNet-style bottleneck blocks with skip connections for channel reduction, MobileNetV3-Small backbone for effective feature extraction and Transformer encoders with multi-head self-attention for robust model sequence. The recognition module handles conjunct consonants, vowel signs and Unicode normalization problems by using a grapheme-based tokenization technique created especially for the complex Bengali letter. For ensuring clean and clear text portion this model also relies on CLAHE along with graysclale at preprocessing period. This hybrid model achieves 84.3% word accuracy and 9.87% Character Error Rate on overall text recognition using a test set of 18,953 word samples. Fixing Unicode Normalization improved accuracy by 11%. The model shows strong performance with 88.4% character accuracy and 92.08% average prediction confidence. The detection part uses YOLOv8n for word localization and for the recognition pipeline applies a hybrid STN-CNN-Transformer structure developed with CTC loss to handle word crops. Finally, this integrated architecture helps the model to gain satisfactory rate of accuracy in case of recognizing both Bengali (85.3%) and English (83.3%) handwriting by keeping the system robust and effective. Also this provides a comprehensive solution for digitising handwritten and printed historical Bengali-English records, offers administration digitisation and academical resources preservations. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Sanjana Zaman | |
| dc.description.statementofresponsibility | Md. Kawser Alam Noman | |
| dc.description.statementofresponsibility | Khandoker Sakib Raiyan | |
| dc.description.statementofresponsibility | Md. Sulaiman Hossain | |
| dc.format.extent | 58 pages | |
| dc.identifier.other | ID 21201051 | |
| dc.identifier.other | ID 22101761 | |
| dc.identifier.other | ID 22201597 | |
| dc.identifier.other | ID 20301010 | |
| dc.identifier.uri | http://hdl.handle.net/10361/27936 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Handwritten text recognition | en_US |
| dc.subject | Bengali-English OCR | en_US |
| dc.subject | Deep learning | en_US |
| dc.subject | MobileNetV3 | en_US |
| dc.subject | Transformer encoder | en_US |
| dc.subject | Bilingual models | en_US |
| dc.subject.lcsh | Optical character recognition. | |
| dc.subject.lcsh | Image processing--Digital techniques. | |
| dc.subject.lcsh | Deep learning (Machine learning)--Mathematical models. | |
| dc.subject.lcsh | Neural networks (Computer science). | |
| dc.subject.lcsh | Writing--Identification--Data processing. | |
| dc.title | Punolikhon: a deep learning approach for handwritten text recognition of Bengali and English documents using a lightweight model | en_US |
| dc.type | Thesis | en_US |