BRAC University Institutional Repository

Research report on Bengla OCR training and testing methods

Show simple item record Hasnat, Md. Abul 2010-10-28T04:03:53Z 2010-10-28T04:03:53Z 2007 2007
dc.description Includes bibliographical references (page 6-7).
dc.description.abstract In this paper we present the training and recognition mechanism of a Hidden Markov Model (HMM) based multi-font Optical Character Recognition (OCR) system for Bengali character. In our approach, the central idea is to separate the HMM model for each segmented character or word. The system uses HTK toolkit for data preparation, model training and recognition. The Features of each trained character are calculated by applying the Discrete Cosine Transform (DCT) to each pixel value of the character image where the image is divided into several frames according to its size. The extracted features of each frame are used as discrete probability distributions which will be given as input parameters to each HMM model. In the case of recognition, a model for each separated character or word is built up using the same approach. This model is given to the HTK toolkit to perform the recognition using the Viterbi Decoding method. The experimental results show significant performance over models using neural network based training and recognition systems. en_US
dc.description.statementofresponsibility Md. Abul Hasnat
dc.format.extent 7 pages
dc.language.iso en en_US
dc.publisher BRAC University en_US
dc.subject Bangla language processing
dc.subject Bangla OCR
dc.title Research report on Bengla OCR training and testing methods en_US
dc.type Technical report en_US
dc.contributor.department Center for Research on Bangla Language Processing (CRBLP), BRAC University

Files in this item

This item appears in the following Collection(s)

Show simple item record

Policy Guidelines

Search BRACU Repository

Advanced Search


My Account