Learning developer style: a stylometric framework for attribution, detection, and profiling
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Azmain, Md. Aquib | |
| dc.contributor.author | Sams, Mohammad Sayed Safi | |
| dc.contributor.author | Nandi, Sudipta | |
| dc.contributor.author | Hami, Nabil Al | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-01-11T05:14:12Z | |
| dc.date.available | 2026-01-11T05:14:12Z | |
| dc.date.copyright | 2025 | |
| dc.date.issued | 2025-10 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 41-42). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025. | en_US |
| dc.description.abstract | Code Stylometry applied to source code is the task of revealing the author or source of a source code file, based on style and structure. Authorship verification has grown to be of great significance with the emergence of the open-source code and issues of plagiarism. This paper gives us a composite research on code stylometry consisting of extensive experimentation on more than 3.2 million Python source code files. To make upstream features well-suited for style tasks, we build a pipeline which extracts both low-level (e.g. token-level, syntax) and high-level (e.g. structural, behavioral) styles. Putting the code snippets in groups according to their author, we will collect a set of more than 44,000 different author styles. We use and test a number of machine learning and deep learning models to predict metadata and notice how stylometry features correlate with metadata. Our pipeline will analyze the coding style of the author in every code and the outcome proves that authorial fingerprinting is attainable with immense precision concerning varying pieces of code. In addition, we evaluated the similarity of coding style among authors by using the stylometric features which were provided by our pipeline. Our model offers a modular system on which future tasks, like plagiarism detection, authorship attribution or codegeneration/ mimicking based on a style can be built upon. This contribution provides a solid methodology, a large stylometric datasets, and high-quality baselines, which will enable future line of research on in-secure-code forensics and author-independent intelligent code-generators. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Mohammad Sayed Safi Sams | |
| dc.description.statementofresponsibility | Sudipta Nandi | |
| dc.description.statementofresponsibility | Nabil Al Hami | |
| dc.format.extent | 50 pages | |
| dc.identifier.other | ID 21301704 | |
| dc.identifier.other | ID 21301534 | |
| dc.identifier.other | ID 21301512 | |
| dc.identifier.uri | http://hdl.handle.net/10361/27419 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Code stylometry | en_US |
| dc.subject | Authorship verification | en_US |
| dc.subject | Authorial fingerprinting | en_US |
| dc.subject | Author-independent codes | en_US |
| dc.subject | Intelligent code generation | en_US |
| dc.subject.lcsh | Computer software--Development. | |
| dc.subject.lcsh | Computer programming. | |
| dc.title | Learning developer style: a stylometric framework for attribution, detection, and profiling | en_US |
| dc.type | Thesis | en_US |