BanglaSarc: a dataset for sarcasm detection

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorApon, Tasnim Sakib
dc.contributor.authorAnan, Ramisa
dc.contributor.authorModhu, Elizabeth Antora
dc.contributor.authorSuter, Arjun
dc.contributor.authorSneha, Ifrit Jamal
dc.contributor.authorAlam, Md. Golam Rabiul
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-08-13T09:37:39Z
dc.date.available2026-08-13T09:37:39Z
dc.date.issued2022-01-01
dc.description.abstractBeing one of the most widely spoken language in the world, the use of Bangla has been increasing in the world of social media as well. Sarcasm is a positive statement or remark with an underlying negative motivation that is extensively employed in today's social media platforms. There has been a significant improvement in sarcasm detection in English over the previous many years, however the situation regarding Bangla sarcasm detection remains unchanged. As a result, it is still difficult to identify sarcasm in bangla, and a lack of high-quality data is a major contributing factor. This article proposes BanglaSarc, a dataset constructed specifically for bangla textual data sarcasm detection. This dataset contains of 5112 comments/status and contents collected from various online social platforms such as Facebook, YouTube, along with a few online blogs. Due to the limited amount of data collection of categorized comments in Bengali, this dataset will aid in the of study identifying sarcasm, recognizing people's emotion, detecting various types of Bengali expressions, and other domains. The dataset is publicly available at https://www.kaggle.com/datasets/sakibapon/banglasarc.
dc.description.versionPublished
dc.format.extent5 Pages
dc.identifier.citationT. S. Apon, R. Anan, E. A. Modhu, A. Suter, I. J. Sneha and M. G. R. Alam, "BanglaSarc: A Dataset for Sarcasm Detection," 2022 IEEE Asia-Pacific Conference on Computer Science and Data Engineering (CSDE), Gold Coast, Australia, 2022, pp. 1-5, doi: 10.1109/CSDE56538.2022.10089322.
dc.identifier.doi10.1109/CSDE56538.2022.10089322
dc.identifier.issn9781665453059
dc.identifier.other2-s2.0-85151081700
dc.identifier.urihttps://hdl.handle.net/10361/29051
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/CSDE56538.2022.10089322
dc.relation.ispartofProceedings of IEEE Asia Pacific Conference on Computer Science and Data Engineering Csde 2022
dc.relation.ispartofseriesProceedings of IEEE Asia Pacific Conference on Computer Science and Data Engineering Csde 2022
dc.relation.urihttps://ieeexplore.ieee.org/document/10089322
dc.subjectBangla Natural Langauge Processing (BNLP)
dc.subjectBangla sarcasm detection
dc.subjectEmotion recognition
dc.subjectData engineering
dc.subject.lcshBengali language--Data processing.
dc.subject.lcshComputational linguistics.
dc.subject.lcshNatural language processing (Computer science).
dc.titleBanglaSarc: a dataset for sarcasm detection
dc.typeConference Proceeding
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.identifier.scopus-author-id57348873600
person.identifier.scopus-author-id57913845400
person.identifier.scopus-author-id57913421800
person.identifier.scopus-author-id57912994600
person.identifier.scopus-author-id57913845500
person.identifier.scopus-author-id26434126600

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
IMG_8345.jpg
Size:
27.35 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: