Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Quality assessment of extracted information from newspaper comment sections using natural language processing

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorSadeque, Farig Yousuf
dc.contributor.authorDeb, Arnob
dc.contributor.authorIslam, Maidul
dc.contributor.authorHossain, Sadab Sifar
dc.contributor.authorAlam, Farjana
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2024-05-09T03:23:08Z
dc.date.available2024-05-09T03:23:08Z
dc.date.copyright©2024
dc.date.issued2024-01
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 39-40).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2024.en_US
dc.description.abstractNewspaper comment section– where readers can leave their opinions– can be an excellent source of information embellishment if used properly. Although there is a risk of fake news and misinformation being spread through the comment section, quality information can also be extracted from these comments that may supplement the original news. From recently performed research, a comment can range between irrelevant to informative– and in our thesis, we would like to identify informative news comments that will further be used to supplement the original news article. We will also identify the level of informativeness of a newspaper comment to figure out whether the task of assigning the Editor’s Pick flag (which is currently done by hand at every large news outlet) with the help of state-of-the-art natural language processing and information extraction techniques. We evaluated the similarity between comments and their respective news articles using transformer models like Sentence BERT. Furthermore, we checked if a comment logically entails using different models, from Simple RNN and LSTM to advanced ones like Roberta and big models like Electra. The final model for Textual Entailment (RoBERTa) task outperformed all the other models by achieving an accuracy of 88.60% and the final model for Textual Similarity (SBERT) task outperformed all the similarity models with an accuracy of 68.49%.en_US
dc.description.degreeBachelor of Science in Computer Science
dc.description.statementofresponsibilityArnob Deb
dc.description.statementofresponsibilityMaidul Islam
dc.description.statementofresponsibilitySadab Sifar Hossain
dc.description.statementofresponsibilityFarjana Alam
dc.format.extent52 pages
dc.identifier.otherID: 23241076
dc.identifier.otherID: 20101309
dc.identifier.otherID: 23341064
dc.identifier.otherID: 20101022
dc.identifier.urihttp://hdl.handle.net/10361/22782
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBrac University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectNatural language processingen_US
dc.subjectInformation extractionen_US
dc.subjectS-BERTen_US
dc.subjectRoBERTaen_US
dc.subjectSimilarityen_US
dc.subject.lcshNatural language processing (Computer science)
dc.titleQuality assessment of extracted information from newspaper comment sections using natural language processingen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
23241076, 20101309, 23341064, 20101022_CSE.pdf
Size:
770.58 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: