An efficient text preprocessing and classification technique for multilingual and transliterated data

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorKhan, Towhid
dc.contributor.authorMallick, David Dew
dc.contributor.authorKhan, Md. Shakiful Islam
dc.contributor.authorHasan, Md Mahadi
dc.contributor.authorAshraf, Faisal Bin
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-20T05:10:11Z
dc.date.available2026-09-20T05:10:11Z
dc.date.issued2022-01-01
dc.description.abstractIndividuals express their opinions on a daily basis through various online platforms in the form of text, comments, and reviews, which is becoming a popular arena for text classification researchers. For a better cause, we aimed to ease the researchers' workload by uncovering the true meaning of their perspective using a text preprocessing method and a text classification algorithm. Our goal was to discover a connection that could work well on multilingual and cross-language data. In this paper, we worked with five different datasets from different languages - English, Bangla, and Banglish (transliterated dataset) - on which we ran our experiment and discovered that byte-mLSTM outperforms with XGB classification model. We achieved more than 90% accuracy on English datasets, 93.43% on Bangla datasets, and 78% on transliterated datasets using mLSTM preprocessing, outperforming the existing performance of traditional approaches on English datasets. For preprocessing, mLSTM performs well across all languages.
dc.description.versionPublished
dc.format.extent366-371
dc.identifier.citationT. Khan, D. D. Mallick, M. S. I. Khan, M. M. Hasan and F. B. Ashraf, "An Efficient Text Preprocessing and Classification Technique for Multilingual and Transliterated Data," 2022 25th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2022, pp. 366-371, doi: 10.1109/ICCIT57492.2022.10054834.
dc.identifier.doi10.1109/ICCIT57492.2022.10054834
dc.identifier.issn9798350346022
dc.identifier.other2-s2.0-85150160864
dc.identifier.urihttps://hdl.handle.net/10361/30056
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/ICCIT57492.2022.10054834
dc.relation.ispartofProceedings of 2022 25th International Conference on Computer and Information Technology Iccit 2022
dc.relation.ispartofseriesProceedings of 2022 25th International Conference on Computer and Information Technology Iccit 2022
dc.relation.urihttps://ieeexplore.ieee.org/document/10054834
dc.subjectComputational modeling
dc.subjectText categorization
dc.subjectPipelines
dc.subjectNeural networks
dc.subjectData processing
dc.subjectData models
dc.subjectClassification algorithms
dc.subjectSentiment analysis
dc.subjectText preprocessing
dc.subjectText classification
dc.subjectTransliterated data
dc.subject.lcshNatural language processing (Computer science).
dc.titleAn efficient text preprocessing and classification technique for multilingual and transliterated data
dc.typeConference Proceeding
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.identifier.scopus-author-id58143396900
person.identifier.scopus-author-id58144164600
person.identifier.scopus-author-id58144321800
person.identifier.scopus-author-id57214844191
person.identifier.scopus-author-id57194202985

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
IMG_8345.jpg
Size:
27.35 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: