Key points are not available for this paper at this time.
Background Large language models (LLMs) have developed rapidly since the release of ChatGPT by OpenAI on November 30, 2022. This vogue comes to mind us “Can we distinguish LLMs-generated texts each other?” “What stylometric features are effective for differentiating LLMs such as fingerprints?” The purpose of this study was to distinguish the 300 Japanese texts (e.g., public comments) generated by six LLMs ChatGPT (GPT5), Claude 3.5, Gemini, Microsoft Copilot, Llama 3.1, and Perplexity. Methods To this end, we explored the effective stylometric features e.g., function-word unigrams, part of speech (POS) bigrams, and phrase patterns using the following analyses: (1) UMAP (Uniform Manifold Approximation and Projection) to visually explore distributional differences among LLMs, (2) Random Forest (RF) and XGBoost with leave-one-out cross-validation to differentiate LLMs, and (3) SHAP (SHapley Additive exPlanations) based on Random Forest to identify effective stylometric features. Results First, UMAP demonstrated the separation of texts among the five LLMs except for Llama 3.1, which displayed substantial overlap with the other five LLMs. Second, RF achieved the highest performance across all stylometric features, with macro F1 scores exceeding 0.95 and reaching 1.00 for several LLMs. The detection performance of XGBoost was lower than that of RF, with the macro F1 scores ranging from 0.88 to 0.94. Finally, SHAP revealed LLM-specific patterns in function-word unigrams, POS bigrams, and phrase patterns. Conclusion These findings indicate that Japanese public comments generated by the six LLMs can be accurately distinguished by focusing on the combination or patterns of stylometric features, suggesting LLM-specific linguistic fingerprints regardless of similarities in the underlying transformer architectures.
Zaitsu et al. (Mon,) studied this question.