Advancements in Speech Recognition: A Systematic Review of Deep Learning Transformer Models, Trends, Innovations, and Future Directions

N/ACitations
Citations of this article
69Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

The transformer is a Deep Learning (DL) model that revolutionized language processing with its self-attention mechanism, enabling parallel processing and improving model efficiency, which dramatically reshaped the landscape of speech recognition technology, based on the ability to efficiently manage the dynamic and context-rich nature of speech. The proposed systematic review in this article critically examines the impact of transformer models on speech recognition, covering the published research over the past seven years (from January 1, 2017, to May 15, 2024). The goals of this article are to synthesize the current knowledge, pinpoint emerging trends, and identify research gaps that could be beneficial for future investigations. From an initial pool of 2,838 publications sourced from leading digital libraries, a rigorous two-step screening process applied to distill high-quality studies relevant to the review criteria. We concentrated our analysis on seven pivotal areas as following: the environmental conditions (neutral versus noisy) addressed by the studies, methods of feature extraction employed, characteristics of the transformer models used, datasets utilized, variations in model efficiency, the influence of noise on model generalizability, and trends in self-supervised learning, ending up with 37 articles to review in this paper. Our findings underscore the transformative potential of transformers in enhancing the accuracy and robustness of speech recognition systems, especially in challenging acoustic environments. In addition, this review highlights areas where more research is needed to make speech recognition even better by using transformer technology.

Cite

CITATION STYLE

APA

Sharrab, Y. O., Attar, H., Eljinini, M. A. H., Al-Omary, Y., & Al-Momani, W. E. (2025). Advancements in Speech Recognition: A Systematic Review of Deep Learning Transformer Models, Trends, Innovations, and Future Directions. IEEE Access. Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1109/ACCESS.2025.3550855

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free