Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Conference Papers (Centre for Research on Bangla Language Processing)

Permanent URI for this collectionhttps://hdl.handle.net/10361/102

Browse

Recent Submissions

Now showing 1 - 20 of 40
  • listelement.badge.dso-type Item ,
    Detecting flames and insults in text
    (BRAC University, 2008-12) Mahmud, Altaf; Ahmed, Kazi Zubair; Khan, Mumit; Center for research on Bangla language processing (CRBLP), BRAC University
    While the internet has become the leading source of information, it is also become the medium for flames, insults and other forms of abusive language, which add nothing to the quality of information available. A human reader can easily distinguish between what is information and what is a flame or any other form of abuse. It is however much more difficult for a language processor to do this automatically. This paper describes a new approach for an automated system to distinguish between information and personal attacks containing insulting or abusive expressions in a given document. In linguistics, insulting or abusive messages are viewed as an extreme subset of the subjective language because of its extreme nature. We create a set of rules to extract the semantic information of a given sentence from the general semantic structure of that sentence to separate information from abusive language.
  • listelement.badge.dso-type Item ,
    Text to speech for Bangla language using festival
    (BRAC University, 2007) Alam, Firoj; Nath, Promila Kanti; Khan, Mumit; Center for Research on Bangla language Processing (CRBLP), BRAC University
    In this paper, we present a Text to Speech (TTS) synthesis system for Bangla language using the open-source Festival TTS engine. Festival is a complete TTS synthesis system, with components supporting front-end processing of the input text, language modeling, and speech synthesis using its signal processing module. The Bangla TTS system proposed here, creates the voice data for festival, and additionally extends festival using its embedded scheme scripting interface to incorporate Bangla language support. Festival is a oncatenative TTS system using diphone or other unit selection speech units. Our TTS implementation uses two different kinds of these concatenative methods supported in Festival: unit selection and multisyn unit selection. The function of a Text-to-Speech system is to convert some language text into its spoken equivalent by a series of modules. These modules, constituting the TTS system are described in detail which is very much helpful for future development. Finally, the quality of synthesized speech is assessed in terms of acceptability and intelligibility.
  • listelement.badge.dso-type Item ,
    Collaborative lexicon development for Bangla
    (BRAC University, 2006) Pavel, Dewan Shahriar Hossain; Sarkar, Asif Iqbal; Shah, Faisal Muhammad; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP), BRAC University
    This paper addresses the issue of building a Bangla lexicon with a collaborative effort through stand alone application and web based interface. The words in the lexicon will be annotated with a combination of tags addressing Parts-of-speech, syntactic, semantic and other grammatical features. Bangla words have been classified into several different parts – of – speech categories including various major word groups and subgroups. This paper aims to provide an integrated user – friendly software interface to the user to annotate a large existing Bangla word set and proposes a mechanism to collaboratively integrate linguists and other interested people into the lexicon build up process. The effort will be a significant progress towards development of a properly annotated lexicon. The outcome of the effort will significantly help in the processes of Morphological Analysis, Automatic grammar Extraction and machine translation for Bangla.
  • listelement.badge.dso-type Item ,
    A proposed automated extraction procedure of Bangla text for corpus creation in unicode
    (BRAC University, 2006) Pavel, Dewan Shahriar Hossain; Sarkar, Asif Iqbal; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP)
    This paper addresses the issue of automated Bangla corpus creation, which will significantly help the processes of lexicon development, morphological analysis, automatic parts of speech detection and automatic grammar extraction and machine translation. The plan is to collect all free Bangla documents on the world wide web and offline documents available and extract all the words in them to make a huge repository of text. This body of text or corpus will be used for several purposes of Bangla language processing after it is converted to Unicode text. The conversion process is also one of the associated and equally important research and development issue. Among several procedures our research focuses on a combination of font and language detection and Unicode conversion of retrieved Bangla text as a solution for automatic Bangla corpus creation and the methodology has been described in the paper.
  • listelement.badge.dso-type Item ,
    A comprehensive roman (English)-to-Bangla transliteration scheme
    (BRAC University, 2006) Naushad UzZaman,; Zaheen, Arnab; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP)
    A transliteration scheme from Roman (English) to Bangla can help increase the use of Bangla in essential and diverse computing areas such as word processing, Internet and mobile communication and information query and retrieval. The Bangla script’s irregular phonetic nature and its large repertoire of consonant clusters (juktakkhors) create a large gap between the pronunciation and the orthography for a given Bangla word. In this paper, we describe a comprehensive Roman (English)-to-Bangla transliteration scheme that is designed to handle the full complexity of the Bangla script. We apply a phonetic encoding scheme to produce intermediate code-strings that facilitate matching pronunciations of input strings and the desired outputs. We also provide graceful degradation to a more conventional direct phonetic mapping in special circumstances. A prototype of our scheme shows significant success in test cases.
  • listelement.badge.dso-type Item ,
    Acoustic analysis of Bangla consonants
    (BRAC University, 2008) Alam, Firoj; Habib, S. M. Murtoza; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP), BRAC University
    This paper describes the acoustic characteristics of Bangla consonants, obtained by analyzing the recordings of male and female voices. First, the duration of each phoneme was identified by averaging both the male and female voice data; then, formant were measured and formant comparison was made for controversial phonemes, which also served to resolve the controversies in the existing phoneme inventories; and finally, a consonant phoneme inventory was designed.
  • listelement.badge.dso-type Item ,
    A comprehensive Bangla spelling checker
    (BRAC University, 2006) Naushad UzZaman,; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP), BRAC University
    We present a comprehensive Bangla spelling checker that improves the quality of suggestions for misspelled words. The complex rules for Bangla spelling presents a significant challenge in producing suggestions for a misspelled word when employing the traditional methods; one must take phonetic similarity into account for suggested alternatives to be reasonably accurate. In Bangla there are several algorithms available for spell checking, however, none of these considers the complex orthographic rules of Bangla. As a result, spelling checker application does not perform well. In this paper, we describe the process of checking the spelling of a Bangla document (i.e. detecting misspelled words, generating suggestions for misspelled word, and ranking the suggestions), compare the methodologies with existing solutions available in the literature, and then propose solutions for each step. Finally, we conclude by showing the performance and evaluation of our proposed solution.
  • listelement.badge.dso-type Item ,
    Segmentation free Bangla OCR using HMM: Training and recognition
    (BRAC University, 2007) Hasnat, Md. Abul; Habib, S. M. Murtoza; Khan, Mumit; Center for Research on Bangla language processing (CRBLP), BRAC University
    The wide area of the application of HMM is in Speech Recognition where each spoken word is considered as a single unit to be recognized from the trained word network. Using this concept some research has been done for character recognition. In this paper, we present the training and recognition mechanism of a Hidden Markov Model (HMM) based multi font supported Optical Character Recognition (OCR) system for Bangla character. In our approach the central idea is separate HMM model for each segmented character or word. We emphasize on word level segmentation and like to consider the single character as a word when the character appears alone after segmentation process is done. The system uses HTK toolkit for data preparation, model training from multiple samples and recognition. Features of each trained character are calculated by applying Discrete Cosine Transform (DCT) to each pixel value of the character image where the image is divided into several frames according to its size. The extracted features of each frame are used as discrete probability distributions that will be given as input parameter to each HMM model. In case of recognition a model for each separated character or word is build up using the same approach. This model is given to the HTK toolkit to perform the recognition using Viterbi Decoding. The experimental result shows significant performance.
  • listelement.badge.dso-type Item ,
    BWN- A software platform for developing Bengali wordnet
    (BRAC University, 2008) Khan, Mumit; Faruqe, Farhana; Center for Research on Bangla Language Processing (CRBLP), BRAC University
    Advanced Natural Language Processing (NLP) applications are increasingly dependent on the availability of linguistic resources, ranging from digital lexica to rich tagged and annotated corpora. While these resources are readily available for digitally advanced languages such as English, these have yet to be developed for widely spoken but digitally immature languages such as Bengali. WordNet is a linguistic resource that can be used in, and for, a variety of applications from a digital dictionary to an automatic machine translator. To create a WordNet for a new language however is a significant challenge, not the least of which is the availability of the lexical data, followed by the software framework to build and manage the data. In this paper, we present BWN, a software framework to build and maintain a Bengali WordNet. We discuss in detail the design and implementation of BWN, concluding with a discussion of how it may be used in future to develop WordNets for other languages as well.
  • listelement.badge.dso-type Item ,
    A decentralised approach to information retrieval for a developing country like Bangladesh
    (BRAC University, 2007) Ali, Hammad; Haque, Nafid; Center for Research on Bangla Language Processing (CRBLP)
    In this paper, we talk about a decentralised information retrieval system which would be suitable for the developing countries that face the problem of limited bandwidth. In this paper we came up with an implementation that uses the existing technology in a novel manner to meet the specific needs of users in such countries. We considered the infrastructural limitations of a developing country like Bangladesh and thus the solution presented here performs just as well in similar situations anywhere else in the world. Our work has been tested for the Bangla language and the same procedure can be applied easily for any other language in any part of the world where there is need for such a system. We had to pick and choose from a set of software packages for one that would best serve our needs. We also had to take into consideration user convenience, for which we had to keep in mind the diverse demographics of people that might have need of such a system. Finally, we came up with the system with all the desired features.
  • listelement.badge.dso-type Item ,
    Integrating Bangla script recognition support in tesseract OCR
    (BRAC University, 2009) Hasnat, Md. Abul; Chowdhury, Muttakinur Rahman; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP)
    Tesseract is considered one of the most accurate free software OCR engines currently available. It was originally developed by Hewlett-Packard from 1985 until 1995, and is currently maintained by Google. At present, Tesseract is capable of only recognizing English, French, Italian, German, Spanish and Dutch. However, it is possible to make Tesseract recognize other scripts if the engine is trained with the requisite data. In this paper, we present a complete methodology to integrate Bangla script recognition support in Tesseract.
  • listelement.badge.dso-type Item ,
    Example based English-Bengali machine translation using wordnet
    (BRAC University, 2009) Salm, Khan Md. Anwarus Salam; Khan, Mumit; Nishino, Tetsuro; Center for Research on Bangla Language Processing (CRBLP)
    In this paper we propose an architecture of English-Bengali Example Based Machine Translation (EBMT) using WordNet. The proposed EBMT system has five steps: 1) Tagging 2) Parsing 3) Prepare the chunks of the sentence using sub-sentential EBMT 4) Using an efficient adapting scheme, match the sentence rule 5) Translate from Source Language (English) to Target Language (Bengali) in the chunk and generate with morphological analysis with the help of WordNet. Using the word senses given by the WordNet we can detect the ambiguity and improve the correctness of translation.
  • listelement.badge.dso-type Item ,
    Development of annotated Bangla speech corpora
    (BRAC University, 2010) Alam, Firoj; Habib, S. M. Murtoza; Sultana, Dil Afroza; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP)
    This paper describes the development procedure of three different Bangla read speech corpora which can be used for phonetic research and developing speech applications. Several criteria were maintained in the corpora development process that includes considering the phonetic and prosodic features during text selection. On the other hand, a specification was maintained in the recording phase as the speaking style is a vital part in speech applications. We also concentrated on proper text normalization, pronunciation, aligning, and labeling. The labeling was done manually – in the present endeavor sentence level labeling (annotation) was completed by maintaining a specification so that it could be expanded in future.
  • listelement.badge.dso-type Item ,
    Rule based automated pronunciation generator
    (BRAC University, 2006) Mosaddeque, Ayesha Binte; UzZaman, Naushad; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP)
    This paper presents a rule based ronunciation generator for Bangla words. It takes a word and finds the pronunciations for the graphemes of the word. A grapheme is a unit in writing that cannot be analyzed into smaller components. Resolving the pronunciation of a polyphone grapheme (i.e. a grapheme that generates more than one phoneme) is the major hurdle that the Automated Pronunciation Generator (APG) encounters. Bangla is partially phonetic in nature, thus we can define rules to handle most of the cases. Besides, up till now we lack a balanced corpus which could be used for a statistical pronunciation generator. As a result, for the time being a rule-based approach towards implementing the APG for Bangla turns out to be efficient.
  • listelement.badge.dso-type Item ,
    N-gram based statistical grammar checker for Bangla and English
    (Center for research on Bangla language processing (CRBLP), BRAC University, 2006) Alam, Md. Jahangir; UzZaman, Naushad; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP)
    This paper describes a statistical grammar checker, which considers the n-gram based analysis of words and POS tags to decide whether the sentence is grammatically correct or not. We employed this technique for both Bangla and English and also described limitation in our approach with possible solutions.
  • listelement.badge.dso-type Item ,
    Minimally segmenting performance Bangla optical character recognition using Kohonen network
    (BRAC University, 2006) Shatil, Adnan Mohammad Shoeb; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP)
    This paper presents a method to use Kohonen neural network based classifier in Bangla Optical Character Recognition (OCR) system, providing much higher performance than the traditional neural network based ones. It describes how Bangla characters are processed, trained and then recognized with the use of a Kohonen network. While there have been significant efforts in using the various types of Artificial ,eural ,etworks (A,,) in optical character recognition, this is the first published account of using a segmentation-free optical character recognition system for Bangla using a Kohonen network. The methodology presented here assumes that the OCR pre-processor has minimally segmented the input words into easily segmentable chunks, and presenting each of these as images to the classification engine described here. The size and the font face used to render the characters are also significant in both training and classification. The images are first converted into grayscale and then to binary images; these images are then scaled to a fit a pre-determined area with a fixed but significant number of pixels. The feature vectors are then extracted from the rectangular pixel map, which in this case is simply a series of 0s and 1s of fixed length. Finally, a Kohonen neural network is chosen for the training and classification process. Although the steps are simple, and the simplest network is chosen for the training and recognition process, the resulting classifier is accurate to better than 98%, depending on the quality of the input images.
  • listelement.badge.dso-type Item ,
    JKimmo: A Multilingual computational mophology frame work for PC-KIMMO
    (BRAC University, 2006) Islam, Md. Zahurul; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP)
    Morphological analysis is of fundamental interest in computational linguistics and language processing. While there are established morphological analyzers for mostly Western and a few other languages using localized interfaces, the same cannot be said for Indic and other less-studied languages for which language processing is just beginning. There are three primary obstacles to computational morphological analysis of these less-studied languages: the generative rules that define the language morphology, the morphological processor, and the computational interface that a linguist can use to experiment with the generative rules. In this paper, we present JKimmo, a multilingual morphological open-source framework that uses the PC-KIMMO two-level morphological processor and provides a localized interface for Bangla morphological analysis. We then apply Jkimmo to Bangla computational morphology, demonstrating both its recognition and generation capabilities. Jkimmo’s internationalization (i18n) frame-work allows easy localization in other languages as well, using a property file for the interface definitions and a transliteration scheme for the analysis.
  • listelement.badge.dso-type Item ,
    Infrastructure for Bangla information retrieval in context of ICT for development
    (BRAC University, 2006) Haque, Nafid; Ali, M Hammad; Abduallah, Matin Saad; Center for Research on Bangla Language Processing (CRBLP)
    In this paper, we talk about developing a search engine and information retrieval system for Bangla. Current work done in this area assumes the use of a particular type of encoding or the availability of particular facilities for the user. We wanted to come up with an implementation that did not require any special features or optimizations in the user end, and would perform just as well in all situations. For this purpose, we picked two case studies to work on in our effort to finding a suitable solution to the problem. While working on these cases, we encountered several problems and had to find our way around these problems. We had to pick and choose from a set of software packages for the one that would best serve our needs. We also had to take into consideration user convenience in using our system, for which we had to keep in mind the diverse demographics of people that might have need for such a system. Finally, we came up with the system, with all the desired features. Some possible future developments also came into mind in the course of our work, which are also mentioned in this paper.
  • listelement.badge.dso-type Item ,
    History (Forward N-Gram) or future (Backward N-Gram)? Which model to consider for N-Gram analysis in Bangla?
    (BRAC University, 2006) Khan, Naira; Habib, Md. Tarek; Alam, Md. Jahangir; Rahman, Rajib; UzZaman, Naushad; Khan, Mumit; Center for Research on Bangla Language Processing, BRAC University
    This paper presents a directional advantage of n-gram modeling in terms of backward or forward n-gram modeling in Bangla. The most commonly used n-gram analysis is predominantly a forward n-gram. However in Bangla it appears that a backward n-gram is repeatedly more successful and yields more grammatical results than a forward n-gram. This paper hypothesizes that the rationale behind this success is the syntactic ordering of constituents in Bangla. Bangla is a head-final specifier-initial language as opposed to English, which is head-initial specifier-initial. Hence in Bangla, the head comes after its argument in a phrase. If an n-gram analysis begins with a head and moves backwards it will stretch to its own argument but if you move for-wards then you'll probably grab the argument of an-other head. As probability of occurrence of heads is higher, probability of depending on a head is also higher and hence a backward n-gram will probably have a greater chance of yielding grammatical results. We carried out several experiments to compare different directional results in different applications with an advantage in the backward direction. This will prove a useful linguistic insight in terms of n-gram based analysis depending upon variations of constituent analysis.
  • listelement.badge.dso-type Item ,
    Developing a computational grammar for Bengali using the HPSG formalism
    (BRAC University, 2006) Khan, Naira; Khan, Mumit; Center for Research on Bangla Language Processing (CRBLP)
    This paper describes the first phase of developing a computational grammar for Bengali using the Head- Driven Phrase Structure Grammar (HPSG) formalism. The HPSG formalism is a highly developed framework that combines computational and psycholinguistic research to provide a tool with which the features particular to a language can be captured and simultaneously provide information about language universals i.e. linguistic phenomena common to diverse languages. The grammar is implemented on a Linguistic Knowledge Building (LKB) system that has both a syntactic and a semantic level. The system is a powerful tool that allows the user to build a parser along with a generator by using the formalism to code in linguistic rules through feature structures and feature unification. This paper provides a set of instructions for using the formulation of HPSG to parse as well as generate grammatical sentences of Bengali. This paper will enable linguists to interpret and formulate features and types with which Bengali or a structurally similar language can be coded in HPSG. It provides a stepping-stone towards the development of a full computational grammar for Bengali which will also provide useful information as a computational description for its sister languages.