Key points are not available for this paper at this time.
Phishing messages have evolved from simple fraud templates into socially engineered texts that exploit anxiety, trust, relational obligation, and culturally embedded norms. In Korean phishing messages, attackers frequently combine institutional authority, family or acquaintance framing, requests for cooperation, and urgency cues to induce concrete victim actions such as money transfer, link clicking, phone contact, app installation, or credential submission. However, prior studies have largely emphasized binary phishing detection while offering limited interpretability regarding how such messages mobilize social and cultural persuasion strategies. This study proposes a culturally aware large language model framework for analyzing social engineering tactics in Korean phishing messages. The framework is built on a multidimensional codebook that represents the message text, phishing label, tactic type, relation type, requested action, cultural lever, and evidence span, enabling structured and explainable analysis beyond simple classification. To operationalize this framework, an OpenChat-based model is fine-tuned with QLoRA to generate structured outputs that jointly predict the phishing status and socially relevant attributes, while evidence-span supervision is incorporated to improve grounding and explanation consistency. The evaluation examines not only phishing-detection performance but also attribute-level prediction accuracy, evidence alignment, parsing reliability, and human-rated usefulness and trustworthiness. By integrating the cultural context, relational framing, and evidence-grounded explanation into LLM-based phishing analysis, this study provides an interpretable analytical framework for Korean phishing messages and an evidence-grounded basis for analyst-supportive phishing triage. On the 82-sample authoritative clean hold-out split, Model D produced error-free label predictions and achieved 0.841 exact-match core and 0.886 span-F1. However, because the evaluation used a single 82-sample internal hold-out split and no independent external corpus, these results should be interpreted as feasibility evidence under leakage-controlled conditions rather than as proof of deployment-level robustness or cross-domain generalization. The main contribution of this study is therefore not improved binary detection over strong lexical baselines, but the structured and evidence-grounded representation of Korean phishing persuasion tactics for analyst-supportive triage.
Lee et al. (Wed,) studied this question.