Parts of speech tagging in Bangla sentences using supervised learning: A performance comparison between viterbi and bidirectional-LSTM models

Loading...
Thumbnail Image

Publisher

Institute of Electrical and Electronics Engineers Inc.

Citation

M. Rumman, A. N. Tasneem and M. G. R. Alam, "Parts of Speech Tagging in Bangla Sentences using Supervised Learning: A Performance Comparison between Viterbi and Bidirectional-LSTM Models," 2021 IEEE International Women in Engineering (WIE) Conference on Electrical and Computer Engineering (WIECON-ECE), Dhaka, Bangladesh, 2021, pp. 9-12, doi: 10.1109/WIECON-ECE54711.2021.9829581.

Abstract

Parts of speech (POS) tagging is a crucial preprocessing step for many Natural Language Processing applications. Though numerous works have been done on English corpus with high accuracy, very few works have been done on Bangla Corpus due to scarcity of resources and the ambiguity of the language. In this paper we have created a POS tagger using Hidden Markov Model(HMM) with Viterbi Algorithm for decoding and a deep learning model called Bidirectional Long-short term memory (BiLSTM). We used similar datasets to compare the performance of the two approaches. It can be inferred from the results that increasing the size of dataset has greater positive impact on the performace of Bi-LSTM model than on the HMM model.

Description

Type

Conference Proceeding