Document template identification and data extraction using machine learning and deep learning approach
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Rhaman, Md. Khalilur | |
| dc.contributor.author | Roy, Kaushik | |
| dc.contributor.author | Islam, Md Fuad | |
| dc.contributor.author | Rimon, Md Minhazul Islam | |
| dc.contributor.author | Mobarak, Tasnim | |
| dc.contributor.author | Priota, Mysha Samiha | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2024-05-26T03:42:05Z | |
| dc.date.available | 2024-05-26T03:42:05Z | |
| dc.date.copyright | ©2024 | |
| dc.date.issued | 2024-01 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 42-43). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2024. | en_US |
| dc.description.abstract | As the world keeps progressing and we continue on our path to a technologically advanced tomorrow, the demand for quick data processing and organization is becoming more and more necessary. People now have access to technology more than ever before. Nowadays, technology allows for the processing and storing of nearly every kind of data. However, procedures requiring paper are still in place and the time-consuming process of moving these data from paper to computers is laborious which reduces work efficiency. Our goal is to make this tedious and time-consuming process fast and efficient, by directly converting the information of the manually checked scripts into digital data. Our research strategy involved gathering information from Brac University examination scripts, digitizing the verified scripts’ data, and then uploading it to a spreadsheet file. The goal of the process is to make Brac University’s grade-processing system quicker, more effective, and less tiresome for the teachers. Three machine learning models and three deep learning models as well as one transfer learning model were utilized for this study. Three common measures were used to evaluate the results which are precision, recall and F1-score. The KNN model showed up to 85% accuracy, whilst SVM showed 87% and SGDClassifier showed 81% accuracy. Meanwhile CNN and YOLOv8 showed 98.6% and 98.8% accuracy respectively. Since YOLOv8 is providing the best accuracy, we will be using this to create an interface that will carry out the complete data transformation process from beginning to end. Starting with capturing the image, processing it to identify the areas from which the data will be collected, and finally extracting the data, in the entire process YOLOv8 is going to be used. In the end, we will obtain precisely extracted data from handwritten exam scripts, which will be arranged in a spreadsheet, digitizing the laborious task of manually inputting each and every grade in a spreadsheet. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science | |
| dc.description.statementofresponsibility | Kaushik Roy | |
| dc.description.statementofresponsibility | Md Fuad Islam | |
| dc.description.statementofresponsibility | Md Minhazul Islam Rimon | |
| dc.description.statementofresponsibility | Tasnim Mobarak | |
| dc.description.statementofresponsibility | Mysha Samiha Priota | |
| dc.format.extent | 55 pages | |
| dc.identifier.other | ID: 20101185 | |
| dc.identifier.other | ID: 20101060 | |
| dc.identifier.other | ID: 20101078 | |
| dc.identifier.other | ID: 20101296 | |
| dc.identifier.other | ID: 20301205 | |
| dc.identifier.uri | http://hdl.handle.net/10361/22915 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | Brac University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | CNN | en_US |
| dc.subject | YOLOv8 | en_US |
| dc.subject | Deep learning model | en_US |
| dc.subject | SGD classifier | en_US |
| dc.subject | SVM | en_US |
| dc.subject | Machine learning | en_US |
| dc.subject | KNN | en_US |
| dc.subject.lcsh | Optical data processing | |
| dc.subject.lcsh | Data structures (Computer science) | |
| dc.subject.lcsh | Neural networks (Computer science) | |
| dc.subject.lcsh | Deep learning (Machine learning) | |
| dc.subject.lcsh | Cognitive learning theory (Deep learning) | |
| dc.title | Document template identification and data extraction using machine learning and deep learning approach | en_US |
| dc.type | Thesis | en_US |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- 20101185, 20101060, 20101078, 20101296, 20301205_CSE.pdf
- Size:
- 769.82 KB
- Format:
- Adobe Portable Document Format
- Description:
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: