Human or Machine? A Survey on Machine-Generated Text Detection

2Citations
Citations of this article
26Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

As generative AI advances rapidly across education, research, medicine, and journalism, machine-generated text (MGT) raises questions about authenticity, ethics, and social impact. To ground this discussion, we conducted a linguistic analysis, covering phonology, morphology, syntax, semantics, lexicon, and pragmatics, which uncovers robust MGT signatures, such as lower perplexity and simpler morphology. We then consolidate the state-of-the-art by reviewing 30 benchmark corpora, totaling over 4 million text samples, and 44 empirical studies, including outcomes from six major shared tasks. Detection approaches are grouped into five broad classes: classical machine learning, deep learning, transformer-based architectures, commercial AI detectors, and statistical tools. Transformer models achieve near-perfect accuracy (≈ 100%), while human evaluators peak at ≈ 77 % accuracy. This survey also highlights common evaluation setups and the core performance measures used to assess model effectiveness. Future MGT detectors can become truly fair, scalable, and effective across languages and domains by expanding corpus diversity and innovating resource-efficient, adversarially robust methods.

Cite

CITATION STYLE

APA

Ahmad, Z., Torres-Ruiz, M., Mahmood, A., Quintero, R., Ameer, I., & Bölücü, N. (2026). Human or Machine? A Survey on Machine-Generated Text Detection. IEEE Access. Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1109/ACCESS.2026.3666781

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free