Authorship verification (AV) remains critical to attribute text provenance and to limit misuse of large language models (LLMs). Existing detectors typically used handcrafted stylometric cues or deep transformer embeddings but lacked robustness to model mimicry and targeted obfuscation and often required large labeled corpora. This work proposes Hybrid Stylometric–Transformer Embedding Framework (HSTEF), a hybrid, explainable AV architecture that fuses multi-level stylometric descriptors with cross-attended transformer embeddings to form a compact verification representation. HSTEF integrated engineered lexical, orthographic and syntactic stylometric vectors, a pre-trained transformer encoder fine-tuned with contrastive and calibration objectives, and a cross-modal fusion module that enabled bidirectional style–semantic interaction while preserving interpretability. The method was evaluated on the PAN CLEF 2025 Voight-Kampff corpus containing human and machine authored texts across genres and including deliberate obfuscations. Experiments used one classical baseline (Linear SVM on TF-IDF, SVM) and one hybrid baseline (DistilBERT + stylometric features) under identical preprocessing and metrics (ROC-AUC, F1, Brier, C@1, and computation). HSTEF produced an absolute ROC-AUC increase of ≈0.07 and an absolute F1 increase of ≈0.07 over the hybrid baseline while improving calibration (Brier reduced ≈47%) at a controlled inference cost increase consistent with forensic deployment scenarios. These results indicate that a Hybrid Stylometric–Transformer Embedding Framework offers a tractable accuracy and robustness trade-off for AV against modern generative AI adversaries.
Journal of Theoretical and Applied Information Technology (Thu,) studied this question.