Ashraf, Faisal BinKhan, TowhidMallick, David DewKhan, Md.Shakiful IslamHasan, Md Mahadi2023-10-172023-10-17©20229/28/2022ID 18201035ID 18201045ID 18201198ID 18201062http://hdl.handle.net/10361/21865Cataloged from PDF version of thesis.Includes bibliographical references (pages 43-44).This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2022.The procedure of eradicating extraneous textual elements and preparing or process- ing the values to be fed into the classifier model is often indicates the concept of text-preprocessing. There are several preprocessing methods, however not all of them are effective when used with cross-language and multilingual datasets. Run- ning a cross-lingual or multilingual dataset through a single pre-processing method and text classification model is rather challenging. What if a technique could be used to better classify data from multilingual and cross lingual datasets? In order to accelerate the process of improving accuracy, we tested various combinations of data pre-processing with text classification models on datasets in Bangla, English, and cross-lingual (Native language written in English letters). We may infer from our experiment that mLSTM functioned effectively for datasets in Bangla and English. Thus, mLSTM can be a helpful preprocessing method for datasets containing a variety of languages.54 pagesenBrac University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.Random forestLogistic regressionTF-IDFSVMXGBmLSTMLSTMInformation retrievalSentiment analysisNLPNatural language processing (Computer science)Computational linguistics--CongressesText classification with an efficient preprocessing technique for cross-language and multilingual dataThesis