Detecting misleading information from Large Language Models responses
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Azmain, Md. Aquib | |
| dc.contributor.advisor | Anwar, Md. Tawhid | |
| dc.contributor.author | Ador, Muntasir Ahmed | |
| dc.contributor.author | Hasan, Fahim | |
| dc.contributor.author | Mahamud, Syed Ashik | |
| dc.contributor.author | Nazmin, Iffat Ara | |
| dc.contributor.author | Muntasir Arin, Md. Mahim | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2025-06-18T05:08:49Z | |
| dc.date.available | 2025-06-18T05:08:49Z | |
| dc.date.copyright | 2025 | |
| dc.date.issued | 2025-02 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 31-34). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025. | en_US |
| dc.description.abstract | The arrival of large language models (LLMs) have been a game-changer in natural language processing (NLP). It revolutionized the way we comprehend and generate content. However, LLMs can hallucinate-that is, contradict the reality or input provided by the user. This is a big problem because these models are being used these days in many diverse industries, including for medical and legal purposes where accuracy is paramount. Hallucinations can damage user trust and lead to the spread of incorrect facts. Although it is not a major issue in ordinary situations, it raises serious concerns in sensitive areas like healthcare and legal advice. Also it can inadvertently become part of the training corpus for future models if not carefully filtered. This creates a feedback loop where errors in one generation of models can propagate and potentially amplify in subsequent iterations. To solve this problem, we have created an all-rounded dataset with questions from SQuAD (Stanford Question Answering Dataset), HotpotQA, and TriviaQA, among other datasets. We will use the state-of-the-art LLM GPT-4o mini to generate answers. Finally, to this end, the paper describes several rules for the annotation of a corresponding dataset, its resulting characteristic properties and the classification quality that can be achieved when using the dataset for fine-tuning different models. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Muntasir Ahmed Ador | |
| dc.description.statementofresponsibility | Fahim Hasan | |
| dc.description.statementofresponsibility | Syed Ashik Mahamud | |
| dc.description.statementofresponsibility | Iffat Ara Nazmin | |
| dc.description.statementofresponsibility | Md. Mahim Muntasir Arin | |
| dc.format.extent | 34 pages | |
| dc.identifier.other | ID: 20101259 | |
| dc.identifier.other | ID: 20201144 | |
| dc.identifier.other | ID: 20301124 | |
| dc.identifier.other | ID: 21101106 | |
| dc.identifier.other | ID:24141109 | |
| dc.identifier.uri | http://hdl.handle.net/10361/26078 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses reports are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Machine learning | en_US |
| dc.subject | Natural language processing | en_US |
| dc.subject | Large language models | en_US |
| dc.subject | AI hallucination | en_US |
| dc.subject | GPT | en_US |
| dc.subject | Transformer | en_US |
| dc.subject | ALBERT | en_US |
| dc.subject.lcsh | Natural language processing (Computer science). | |
| dc.subject.lcsh | Artificial intelligence. | |
| dc.title | Detecting misleading information from Large Language Models responses | en_US |
| dc.type | Thesis | en_US |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- 20101259, 20201144, 20301124, 21101106, 24141109_CSE.pdf
- Size:
- 805.25 KB
- Format:
- Adobe Portable Document Format
- Description:
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: