BRAC University Institutional Repository

Rule based segmentation of lower modifiers in complex Bangla scripts

DSpace/Manakin Repository

Show simple item record

dc.contributor.author Hasnat, Md. Abul
dc.contributor.author Khan, Mumit
dc.date.accessioned 2010-10-05T06:16:48Z
dc.date.available 2010-10-05T06:16:48Z
dc.date.copyright 2009
dc.date.issued 2009
dc.identifier.uri http://hdl.handle.net/10361/338
dc.description Includes bibliographical references (page 5).
dc.description.abstract Segmentation is the most challenging part of Bangla optical character recognition (OCR). To solve the problems of joining errors, several algorithms have been proposed in the literature, with varying degrees of accuracy. The selection of the lower modifier container units and the subsequent extraction of the modifiers from the core unit during segmentation have not been studied extensively. We present a dissection based lower modifier segmentation method which solves the problem of segmenting lower modifiers under a wide range of document images. A key goal in our methodology is to avoid over-segmentation of the units that do not actually contain any lower modifier, leading to unacceptably high error rates during segmentation. Our methodology consists of four tasks: we first identify the lower modifier separator line using character height information, and then select the primary lower modifier containers; we filter this set to eliminate the units/characters that do not actually contain any lower modifier; we then extract the lower modifier unit using the features of the core units and the lower modifiers; the final step consists of a set of empirical rules, aided by dictionary lookups, to eliminate most of the errors, resulting in an accuracy of 99.6%. en_US
dc.description.statementofresponsibility Md. Abul Hasnat
dc.description.statementofresponsibility Mumit Khan
dc.format.extent 8 pages
dc.language.iso en en_US
dc.publisher BRAC University en_US
dc.title Rule based segmentation of lower modifiers in complex Bangla scripts en_US
dc.type Article en_US
dc.contributor.department Center for Research on Bangla Language Processing (CRBLP), BRAC University


Files in this item

Files Size Format View
Rule based segm ... x Bangla scripts, 2009.pdf 226.3Kb PDF View/Open or Preview

This item appears in the following Collection(s)

Show simple item record

Policy Guidelines

Search DSpace


Browse

My Account