Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Text classification with an efficient preprocessing technique for cross-language and multilingual data

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorAshraf, Faisal Bin
dc.contributor.authorKhan, Towhid
dc.contributor.authorMallick, David Dew
dc.contributor.authorKhan, Md.Shakiful Islam
dc.contributor.authorHasan, Md Mahadi
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2023-10-17T08:43:07Z
dc.date.available2023-10-17T08:43:07Z
dc.date.copyright©2022
dc.date.issued9/28/2022
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 43-44).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2022.en_US
dc.description.abstractThe procedure of eradicating extraneous textual elements and preparing or process- ing the values to be fed into the classifier model is often indicates the concept of text-preprocessing. There are several preprocessing methods, however not all of them are effective when used with cross-language and multilingual datasets. Run- ning a cross-lingual or multilingual dataset through a single pre-processing method and text classification model is rather challenging. What if a technique could be used to better classify data from multilingual and cross lingual datasets? In order to accelerate the process of improving accuracy, we tested various combinations of data pre-processing with text classification models on datasets in Bangla, English, and cross-lingual (Native language written in English letters). We may infer from our experiment that mLSTM functioned effectively for datasets in Bangla and English. Thus, mLSTM can be a helpful preprocessing method for datasets containing a variety of languages.en_US
dc.description.degreeBachelor of Science in Computer Science
dc.description.statementofresponsibilityTowhid Khan
dc.description.statementofresponsibilityDavid Dew Mallick
dc.description.statementofresponsibilityMd.Shakiful Islam Khan
dc.description.statementofresponsibilityMd Mahadi Hasan
dc.format.extent54 pages
dc.identifier.otherID 18201035
dc.identifier.otherID 18201045
dc.identifier.otherID 18201198
dc.identifier.otherID 18201062
dc.identifier.urihttp://hdl.handle.net/10361/21865
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBrac University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectRandom foresten_US
dc.subjectLogistic regressionen_US
dc.subjectTF-IDFen_US
dc.subjectSVMen_US
dc.subjectXGBen_US
dc.subjectmLSTMen_US
dc.subjectLSTMen_US
dc.subjectInformation retrievalen_US
dc.subjectSentiment analysisen_US
dc.subjectNLPen_US
dc.subject.lcshNatural language processing (Computer science)
dc.subject.lcshComputational linguistics--Congresses
dc.titleText classification with an efficient preprocessing technique for cross-language and multilingual dataen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
18201035, 18201045, 18201198, 18201062_CSE.pdf
Size:
8.12 MB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: