Robust detection of AI-generated text using stylometric-semantic modeling under paraphrasing and adversarial rewriting
| bracu.embargo.enddate | ||
| bracu.type.group | Research Publications | |
| datacite.rights | Metadata Only | |
| dc.contributor.author | Tasnim, Sibgatullah | |
| dc.contributor.author | Khondoker, Ahmed Zarif | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-08-10T04:41:35Z | |
| dc.date.available | 2026-08-10T04:41:35Z | |
| dc.date.issued | 2026-06-11 | |
| dc.description.abstract | The widespread usage of large language models like ChatGPT, Claude and Gemini has made it harder to distinguish between human and AI-generated writing. This research describes a fully reproducible and complete AI-generated text recognition pipeline that uses stylistic and semantic (embedding-based) properties to detect text well. The public dataset, Human vs. LLM Text Corpus was used for this research. This study can examine its proposed detection methods at a realistic sample size because this dataset has 25 times more samples than previous studies. The research analyzes adversarial robustness against paraphrase assaults and evaluates TF-IDF-based classifiers, sentence-embedding models and hybrid fusion architectures. The TF-IDF baseline outperformed embedding-based approaches with an accuracy of 83.14% and a ROC-AUC of 91.87% on the complete unbalanced dataset. Using a balanced dataset, a hybrid LightGBM model obtained 80.98% accuracy and showed strong robustness, with just a 2.52% loss in F1-score after paraphrase attacks. Document-level stylistic features dominated the model's decision-making process, with word count being the most discriminative variable (importance = 646), despite accounting for less than 1% of all features. This study presents a clear, large-scale benchmark for detecting AI-generated content and shows that feature interpretability and structural indications are needed for accurate AI authorship verification. | |
| dc.description.version | Published | |
| dc.format.extent | 6 pages | |
| dc.identifier.citation | S. Tasnim and A. Z. Khondoker, "Robust Detection of AI-Generated Text Using Stylometric-Semantic Modeling Under Paraphrasing and Adversarial Rewriting," 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN), Chittagong, Bangladesh, 2026, pp. 1-6, doi: 10.1109/QPAIN69676.2026.11545549. | |
| dc.identifier.doi | 10.1109/QPAIN69676.2026.11545549 | |
| dc.identifier.issn | 979-833154990-9 | |
| dc.identifier.uri | https://hdl.handle.net/10361/28856 | |
| dc.language.iso | en_US | |
| dc.publisher | Institute of Electrical and Electronics Engineers Inc. | |
| dc.relation.uri | https://ieeexplore.ieee.org/document/11545549 | |
| dc.subject | Adversarial text augmentation | |
| dc.subject | AI-generated text detection | |
| dc.subject | Hybrid models | |
| dc.subject | Paraphrase robustness | |
| dc.subject | Sentence embeddings | |
| dc.subject | Stylometric | |
| dc.subject.lcsh | Artificial intelligence. | |
| dc.subject.lcsh | Digital media--Security measures. | |
| dc.subject.lcsh | Pattern recognition systems. | |
| dc.subject.lcsh | Machine learning. | |
| dc.title | Robust detection of AI-generated text using stylometric-semantic modeling under paraphrasing and adversarial rewriting | |
| dc.type | Conference Proceedings |