Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Machine learning dataset for criminal law suggestions using case studies within the context of Bangladesh

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorMostakim, Moin
dc.contributor.authorTasfin, Ahnaf Arif
dc.contributor.authorIslam, Wasif
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2025-05-22T03:11:53Z
dc.date.available2025-05-22T03:11:53Z
dc.date.copyright2025
dc.date.issued2025-02
dc.descriptionCataloged from PDF version of internship report.
dc.descriptionIncludes bibliographical references (pages 37-38).
dc.descriptionThis internship report is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.en_US
dc.description.abstractWith the growth of Language Modeling and the upcoming natural language processing assisted tools which aim for text generation, can someday render bureaucracy mean ingless while also decreasing human workload. To bring about that day, we need datasets which are able to train those models. Especially in the case of Bangladesh where there are very few datasets based on Bangladesh’s legal case studies. In this paper, we have created a Multi-Label and Binary classification dataset through data augmentation using verified case studies from the Manupatra database with the criminal law subject within the context of Bangladesh and have tested them with a few models such as DistilBERT, BERT, GPT-2, GPT-3, XLNet and Mamba to classify the acts involved and court in a case study for a 2000 data split and 3000 data split. So far, we have collected 3000 criminal case studies for our augmented dataset. Experimental results showed that out of all the models, Mamba performed the best while GPT-2 came up with the worst results. DistilBERT showed almost similar results to BERT and XLNet despite their computational differences during the benchmarking process of our augmented dataset.
dc.description.degreeBachelor of Science in Computer Science
dc.description.statementofresponsibilityAhnaf Arif Tasfin
dc.description.statementofresponsibilityWasif Islam
dc.format.extent38 pages
dc.identifier.otherID 21101156
dc.identifier.otherID 21101199
dc.identifier.urihttp://hdl.handle.net/10361/25974
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectMachine learningen_US
dc.subjectNatural language processingen_US
dc.subjectBidirectional encoder representationsen_US
dc.subjectNeural networksen_US
dc.subjectManupatraen_US
dc.subject.lcshMachine learning.
dc.subject.lcshData learning.
dc.subject.lcshNatural language processing (Computer science).
dc.subject.lcshHuman-computer interaction.
dc.titleMachine learning dataset for criminal law suggestions using case studies within the context of Bangladeshen_US
dc.typeInternship Reporten_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
21101156,21101199_CSE.pdf
Size:
261.7 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: