An efficient text preprocessing and classification technique for multilingual and transliterated data
| bracu.type.group | Research Publications | |
| datacite.rights | Metadata Only | |
| dc.contributor.author | Khan, Towhid | |
| dc.contributor.author | Mallick, David Dew | |
| dc.contributor.author | Khan, Md. Shakiful Islam | |
| dc.contributor.author | Hasan, Md Mahadi | |
| dc.contributor.author | Ashraf, Faisal Bin | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-09-20T05:10:11Z | |
| dc.date.available | 2026-09-20T05:10:11Z | |
| dc.date.issued | 2022-01-01 | |
| dc.description.abstract | Individuals express their opinions on a daily basis through various online platforms in the form of text, comments, and reviews, which is becoming a popular arena for text classification researchers. For a better cause, we aimed to ease the researchers' workload by uncovering the true meaning of their perspective using a text preprocessing method and a text classification algorithm. Our goal was to discover a connection that could work well on multilingual and cross-language data. In this paper, we worked with five different datasets from different languages - English, Bangla, and Banglish (transliterated dataset) - on which we ran our experiment and discovered that byte-mLSTM outperforms with XGB classification model. We achieved more than 90% accuracy on English datasets, 93.43% on Bangla datasets, and 78% on transliterated datasets using mLSTM preprocessing, outperforming the existing performance of traditional approaches on English datasets. For preprocessing, mLSTM performs well across all languages. | |
| dc.description.version | Published | |
| dc.format.extent | 366-371 | |
| dc.identifier.citation | T. Khan, D. D. Mallick, M. S. I. Khan, M. M. Hasan and F. B. Ashraf, "An Efficient Text Preprocessing and Classification Technique for Multilingual and Transliterated Data," 2022 25th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2022, pp. 366-371, doi: 10.1109/ICCIT57492.2022.10054834. | |
| dc.identifier.doi | 10.1109/ICCIT57492.2022.10054834 | |
| dc.identifier.issn | 9798350346022 | |
| dc.identifier.other | 2-s2.0-85150160864 | |
| dc.identifier.uri | https://hdl.handle.net/10361/30056 | |
| dc.language.iso | en_US | |
| dc.publisher | Institute of Electrical and Electronics Engineers Inc. | |
| dc.relation.hasversion | 10.1109/ICCIT57492.2022.10054834 | |
| dc.relation.ispartof | Proceedings of 2022 25th International Conference on Computer and Information Technology Iccit 2022 | |
| dc.relation.ispartofseries | Proceedings of 2022 25th International Conference on Computer and Information Technology Iccit 2022 | |
| dc.relation.uri | https://ieeexplore.ieee.org/document/10054834 | |
| dc.subject | Computational modeling | |
| dc.subject | Text categorization | |
| dc.subject | Pipelines | |
| dc.subject | Neural networks | |
| dc.subject | Data processing | |
| dc.subject | Data models | |
| dc.subject | Classification algorithms | |
| dc.subject | Sentiment analysis | |
| dc.subject | Text preprocessing | |
| dc.subject | Text classification | |
| dc.subject | Transliterated data | |
| dc.subject.lcsh | Natural language processing (Computer science). | |
| dc.title | An efficient text preprocessing and classification technique for multilingual and transliterated data | |
| dc.type | Conference Proceeding | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.identifier.scopus-author-id | 58143396900 | |
| person.identifier.scopus-author-id | 58144164600 | |
| person.identifier.scopus-author-id | 58144321800 | |
| person.identifier.scopus-author-id | 57214844191 | |
| person.identifier.scopus-author-id | 57194202985 |