Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Sentiment analysis for Bangla microblog posts

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.authorShaika, Chowdhury
dc.contributor.authorChowdhury, Wasifa
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2014-01-29T06:58:08Z
dc.date.available2014-01-29T06:58:08Z
dc.date.issued2014-01
dc.descriptionCataloged from PDF version of thesis report.
dc.descriptionIncludes bibliographical references (page 47).
dc.descriptionThis thesis report is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2014.en_US
dc.description.abstractSentiment analysis has received great attention recently due to the huge amount of user-generated information on the microblogging sites, such as Twitter [1], which are utilized for many applications like product review mining and making future predictions of events such as predicting election results. Much of the research work on sentiment analysis has been applied to the English language, but construction of resources and tools for sentiment analysis in languages other than English is a growing need since the microblog posts are not just posted in English, but in other languages as well. Work on Bangla (or Bengali language) is necessary as it is one of the most spoken languages, ranked seventh in the world [13]. In this paper, we aim to automatically extract the sentiments or opinions conveyed by users from Bangla microblog posts and then identify the overall polarity of texts as either negative or positive. We use a semi-supervised bootstrapping approach for the development of the training corpus which avoids the need for labor intensive manual annotation. For classification, we use Support Vector Machines (SVM) and Maximum Entropy (MaxEnt) and do a comparative analysis on the performance of these two machine learning algorithms by experimenting with a combination of various sets of features. We also construct a Twitter-specific Bangla sentiment lexicon, which is utilized for the rule-based classifier and as a binary feature in the classifiers used. For our work, we choose Twitter as the microblogging site as it is one of the most popular microblogging platforms in the world.en_US
dc.description.degreeBachelor of Science in Computer Science and Engineering
dc.description.statementofresponsibilityChowdhury, Shaika
dc.description.statementofresponsibilityChowdhury, Wasifa
dc.format.extent47 pages
dc.identifier.otherID 10101037
dc.identifier.otherID 10101038
dc.identifier.urihttp://hdl.handle.net/10361/2902
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University thesis are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectComputer science and engineering
dc.titleSentiment analysis for Bangla microblog postsen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
10101037 & 10101038.pdf
Size:
1.92 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: