Robust detection of AI-generated text using stylometric-semantic modeling under paraphrasing and adversarial rewriting

bracu.embargo.enddate
bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorTasnim, Sibgatullah
dc.contributor.authorKhondoker, Ahmed Zarif
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-08-10T04:41:35Z
dc.date.available2026-08-10T04:41:35Z
dc.date.issued2026-06-11
dc.description.abstractThe widespread usage of large language models like ChatGPT, Claude and Gemini has made it harder to distinguish between human and AI-generated writing. This research describes a fully reproducible and complete AI-generated text recognition pipeline that uses stylistic and semantic (embedding-based) properties to detect text well. The public dataset, Human vs. LLM Text Corpus was used for this research. This study can examine its proposed detection methods at a realistic sample size because this dataset has 25 times more samples than previous studies. The research analyzes adversarial robustness against paraphrase assaults and evaluates TF-IDF-based classifiers, sentence-embedding models and hybrid fusion architectures. The TF-IDF baseline outperformed embedding-based approaches with an accuracy of 83.14% and a ROC-AUC of 91.87% on the complete unbalanced dataset. Using a balanced dataset, a hybrid LightGBM model obtained 80.98% accuracy and showed strong robustness, with just a 2.52% loss in F1-score after paraphrase attacks. Document-level stylistic features dominated the model's decision-making process, with word count being the most discriminative variable (importance = 646), despite accounting for less than 1% of all features. This study presents a clear, large-scale benchmark for detecting AI-generated content and shows that feature interpretability and structural indications are needed for accurate AI authorship verification.
dc.description.versionPublished
dc.format.extent6 pages
dc.identifier.citationS. Tasnim and A. Z. Khondoker, "Robust Detection of AI-Generated Text Using Stylometric-Semantic Modeling Under Paraphrasing and Adversarial Rewriting," 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN), Chittagong, Bangladesh, 2026, pp. 1-6, doi: 10.1109/QPAIN69676.2026.11545549.
dc.identifier.doi10.1109/QPAIN69676.2026.11545549
dc.identifier.issn979-833154990-9
dc.identifier.urihttps://hdl.handle.net/10361/28856
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.urihttps://ieeexplore.ieee.org/document/11545549
dc.subjectAdversarial text augmentation
dc.subjectAI-generated text detection
dc.subjectHybrid models
dc.subjectParaphrase robustness
dc.subjectSentence embeddings
dc.subjectStylometric
dc.subject.lcshArtificial intelligence.
dc.subject.lcshDigital media--Security measures.
dc.subject.lcshPattern recognition systems.
dc.subject.lcshMachine learning.
dc.titleRobust detection of AI-generated text using stylometric-semantic modeling under paraphrasing and adversarial rewriting
dc.typeConference Proceedings

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Demo.jpg
Size:
27.28 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: