Zero-Shot Classification of Illicit Dark Web Content with Commercial LLMs: A Comparative Study on Accuracy, Human Consistency, and Inter-Model Agreement

2Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

This study evaluates the zero-shot classification performance of eight commercial large language models (LLMs), GPT-4o, GPT-4o Mini, GPT-3.5 Turbo, Claude 3.5 Haiku, Gemini 2.0 Flash, DeepSeek Chat, DeepSeek Reasoner, and Grok, using the CoDA dataset (n = 10,000 Dark Web documents). Results show strong macro-F1 scores across models, led by DeepSeek Chat (0.870), Grok (0.868), and Gemini 2.0 Flash (0.861). Alignment with human annotations was high, with Cohen’s Kappa above 0.840 for top models and Krippendorff’s Alpha reaching 0.871. Inter-model consistency was highest between Claude 3.5 Haiku and GPT-4o (κ = 0.911), followed by DeepSeek Chat and Grok (κ = 0.909), and Claude 3.5 Haiku with Gemini 2.0 Flash (κ = 0.907). These findings confirm that state-of-the-art LLMs can reliably classify illicit content under zero-shot conditions, though performance varies by model and category.

Cite

CITATION STYLE

APA

Prado-Sánchez, V. P., Domínguez-Díaz, A., De-Marcos, L., & Martínez-Herráiz, J. J. (2025). Zero-Shot Classification of Illicit Dark Web Content with Commercial LLMs: A Comparative Study on Accuracy, Human Consistency, and Inter-Model Agreement. Electronics (Switzerland), 14(20). https://doi.org/10.3390/electronics14204101

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free