PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 3, 20260 citations

Machine Learning-Based Detection of AI-Generated Text via Stylistic and Statistical Feature Modeling

View Full Paper
KSKarla SchäferMSM. Steinebach

Key Points

  • AI-generated text is detectable using machine learning models with an F1-score over 90%.
  • The analysis involved correlating 220 stylistic and statistical features of texts from multiple LLMs.
  • Feature modeling includes examining high kuperman age for AI texts versus lexical richness in human writing.
  • Results highlight the potential for developing robust tools against disinformation generated by AI systems.

Abstract

Through the advances of large-language models (LLMs) AI- generated text can be created with ease. But, these tools can also pose a threat, e.g. through the creation of disinformation. In this work, we analysed texts generated by three LLMs: GPT-3.5, LLaMA3, and Qwen from the CUDRT dataset. We extracted 220 stylistic and statistical features of human and AI-generated text using the LFTK library. First, we analysed the features using the pearson correlation. Second, we trained five machine learning models and tested the classifiers on detecting completely AI-generated, polished, rewritten texts, and summaries created by AI. We calculated an F1-score of 90%+ for the text generated entirely by AI, depending on the LLM used. We found that AI-generated texts, independent of LLM, can be identified through a high kuperman age, i.e. high word complexity, whereby human-written texts are written with higher lexical variation and richness. We provide an explanation for the classification results and a comparison with RoBERTa (fine-tuned).

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Schäfer et al. (2025) studied this question.

synapsesocial.com/papers/69a7665fbadf0bb9e87dcc52https://doi.org/10.1109/trustcom66490.2025.00134
Ask AI
Helpful
Bookmark
Share
View Full Paper