Abstract
As generative AI advances rapidly across education, research, medicine, and journalism, machine-generated text (MGT) raises questions about authenticity, ethics, and social impact. To ground this discussion, we conducted a linguistic analysis, covering phonology, morphology, syntax, semantics, lexicon, and pragmatics, which uncovers robust MGT signatures, such as lower perplexity and simpler morphology. We then consolidate the state-of-the-art by reviewing 30 benchmark corpora, totaling over 4 million text samples, and 44 empirical studies, including outcomes from six major shared tasks. Detection approaches are grouped into five broad classes: classical machine learning, deep learning, transformer-based architectures, commercial AI detectors, and statistical tools. Transformer models achieve near-perfect accuracy (≈ 100%), while human evaluators peak at ≈ 77 % accuracy. This survey also highlights common evaluation setups and the core performance measures used to assess model effectiveness. Future MGT detectors can become truly fair, scalable, and effective across languages and domains by expanding corpus diversity and innovating resource-efficient, adversarially robust methods.
Author supplied keywords
Cite
CITATION STYLE
Ahmad, Z., Torres-Ruiz, M., Mahmood, A., Quintero, R., Ameer, I., & Bölücü, N. (2026). Human or Machine? A Survey on Machine-Generated Text Detection. IEEE Access. Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1109/ACCESS.2026.3666781
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.