Khan, TowhidMallick, David DewKhan, Md. Shakiful IslamHasan, Md MahadiAshraf, Faisal Bin2026-09-202026-09-202022-01-01T. Khan, D. D. Mallick, M. S. I. Khan, M. M. Hasan and F. B. Ashraf, "An Efficient Text Preprocessing and Classification Technique for Multilingual and Transliterated Data," 2022 25th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2022, pp. 366-371, doi: 10.1109/ICCIT57492.2022.10054834.97983503460222-s2.0-85150160864https://hdl.handle.net/10361/30056Individuals express their opinions on a daily basis through various online platforms in the form of text, comments, and reviews, which is becoming a popular arena for text classification researchers. For a better cause, we aimed to ease the researchers' workload by uncovering the true meaning of their perspective using a text preprocessing method and a text classification algorithm. Our goal was to discover a connection that could work well on multilingual and cross-language data. In this paper, we worked with five different datasets from different languages - English, Bangla, and Banglish (transliterated dataset) - on which we ran our experiment and discovered that byte-mLSTM outperforms with XGB classification model. We achieved more than 90% accuracy on English datasets, 93.43% on Bangla datasets, and 78% on transliterated datasets using mLSTM preprocessing, outperforming the existing performance of traditional approaches on English datasets. For preprocessing, mLSTM performs well across all languages.366-371en-USComputational modelingText categorizationPipelinesNeural networksData processingData modelsClassification algorithmsSentiment analysisText preprocessingText classificationTransliterated dataNatural language processing (Computer science).An efficient text preprocessing and classification technique for multilingual and transliterated dataConference Proceeding10.1109/ICCIT57492.2022.10054834