This study evaluates the zero-shot classification performance of eight commercial large language models (LLMs), GPT-4o, GPT-4o Mini, GPT-3.5 Turbo, Claude 3.5 Haiku, Gemini 2.0 Flash, DeepSeek Chat, DeepSeek Reasoner, and Grok, using the CoDA dataset (n = 10,000 Dark Web documents). Results show strong macro-F1 scores across models, led by DeepSeek Chat (0.870), Grok (0.868), and Gemini 2.0 Flash (0.861). Alignment with human annotations was high, with Cohen’s Kappa above 0.840 for top models and Krippendorff’s Alpha reaching 0.871. Inter-model consistency was highest between Claude 3.5 Haiku and GPT-4o (κ = 0.911), followed by DeepSeek Chat and Grok (κ = 0.909), and Claude 3.5 Haiku with Gemini 2.0 Flash (κ = 0.907). These findings confirm that state-of-the-art LLMs can reliably classify illicit content under zero-shot conditions, though performance varies by model and category.
Prado-Sánchez et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: