Abstract
Automatic terminology or term extraction (ATE) is a Natural Language Processing (NLP) task intended to automatically identify specialized terms present in domain-specific corpora. As units of knowledge in a specific field of expertise, extracted terms are not only beneficial for several terminographical tasks, but also support and improve several complex downstream tasks, e.g., information retrieval, machine translation, topic detection, and sentiment analysis. ATE systems and datasets annotated for the task at hand have been studied and developed for decades, but more recent approaches have increasingly involved novel neural systems. Despite a large amount of new research on ATE tasks, systematic survey studies covering novel neural approaches are lacking, especially when it comes to the usage of large-scale language models (LLMs). We present a comprehensive survey of neural approaches to ATE, focusing on transformer-based neural models and the recent generative approaches based on LLMs. The study also compares these systems and previous ML-based approaches, which employed feature engineering and non-neural supervised learning algorithms.
Author supplied keywords
Cite
CITATION STYLE
Tran, H. T. H., Martinc, M., Caporusso, J., Delaunay, J., Doucet, A., & Pollak, S. (2026). Recent Advances in Automatic Term Extraction: A Comprehensive Survey. ACM Computing Surveys, 58(9). https://doi.org/10.1145/3787584
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.