BRAC University Institutional Repository

Rule based segmentation of lower modifiers in complex Bangla scripts

Show simple item record Hasnat, Md. Abul Khan, Mumit 2010-10-05T06:16:48Z 2010-10-05T06:16:48Z 2009 2009
dc.description Includes bibliographical references (page 5).
dc.description.abstract Segmentation is the most challenging part of Bangla optical character recognition (OCR). To solve the problems of joining errors, several algorithms have been proposed in the literature, with varying degrees of accuracy. The selection of the lower modifier container units and the subsequent extraction of the modifiers from the core unit during segmentation have not been studied extensively. We present a dissection based lower modifier segmentation method which solves the problem of segmenting lower modifiers under a wide range of document images. A key goal in our methodology is to avoid over-segmentation of the units that do not actually contain any lower modifier, leading to unacceptably high error rates during segmentation. Our methodology consists of four tasks: we first identify the lower modifier separator line using character height information, and then select the primary lower modifier containers; we filter this set to eliminate the units/characters that do not actually contain any lower modifier; we then extract the lower modifier unit using the features of the core units and the lower modifiers; the final step consists of a set of empirical rules, aided by dictionary lookups, to eliminate most of the errors, resulting in an accuracy of 99.6%. en_US
dc.description.statementofresponsibility Md. Abul Hasnat
dc.description.statementofresponsibility Mumit Khan
dc.format.extent 8 pages
dc.language.iso en en_US
dc.publisher BRAC University en_US
dc.title Rule based segmentation of lower modifiers in complex Bangla scripts en_US
dc.type Article en_US
dc.contributor.department Center for Research on Bangla Language Processing (CRBLP), BRAC University

Files in this item

This item appears in the following Collection(s)

Show simple item record

Policy Guidelines

Search BRACU Repository

Advanced Search


My Account