The manual classification of diplomatic documents by security sensitivity level is labor-intensive and inconsistent at scale. This study proposes SACHN-DeBERTa-v3-Large, a hybrid classification architecture integrating a DeBERTa-v3-Large backbone with a Security-Aware Gate, Prototype Classification Layer, and Supervised Contrastive Projection Head. A rule-based preprocessing pipeline removes inline classification markers prior to training, ensuring that models learn semantic content rather than surface-form annotation artifacts. Two diplomatic corpora are evaluated: the WikiLeaks Cable Classifier (9005 documents, three classes) and a novel Foreign Relations of the United States (FRUS) dataset (24,706 documents, four classes) constructed for this study. A controlled multi-LLM comparative study evaluates LLaMA-3-8B-Instruct, Qwen2.5-14B-Instruct, and Mistral-7B-Instruct-v0.3 under identical QLoRA fine-tuning conditions. SACHN-DeBERTa-v3-Large achieves 96.12% accuracy and 92.57% macro F1 on WikiLeaks, and 92.11% accuracy and 91.02% macro F1 on FRUS, surpassing all baselines by at least 10 percentage points in accuracy and 26 points in macro F1 (McNemar’s test, p < 0.001). Under the primary QLoRA protocol (r = 16), no evaluated LLM exceeded 57.16% accuracy; extended experiments with r = 64 and 8-bit INT8 quantisation yielded a best result of 64.48% (Mistral-7B-Instruct, WikiLeaks), confirming that the performance gap relative to SACHN-DeBERTa-v3-Large remains structural. Post-hoc SHAP and LIME analyses confirm that predictions are grounded in domain-specific semantic content, validating the classification marker removal methodology.
Sarıçiçek et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: