Assessment of L2 Spanish pronunciation accuracy via Automatic Speech Recognition

  • Sarymsakova A
  • Martín Rodilla P
N/ACitations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

This study contributes to the evaluation of non-native Spanish speakers’ acoustic production using Artificial Intelligence (AI) tools, specifically Automatic Speech Recognition (ASR) models. In order to determine whether leading ASR models can provide adequate feedback on L2 Spanish pronunciation, we evaluated four models (Wav2Vec, Whisper-large-v2, Whisper-large-v3, and SeamlessM4T) using datasets of non-native Spanish speakers with English, Russian, and German as L1s. Based on a Word Error Rate and Character Error Rate evaluation framework, Whisper-large-v3 and SeamlessM4T demonstrated the highest accuracy for non-native speech recognition. A qualitative and phonetic error analysis revealed that these models struggle when vowel formant boundaries of L2 speakers exceed those of standard Spanish or when voiceless consonants are influenced by phonetic assimilation processes. Additionally, we identified gender bias, with models performing better on female speech than male speech, and substitution errors as the most frequent error type. In conclusion, while ASR models like Whisper-large-v3 and SeamlessM4T perform adequately, an accurate pronunciation assessment for L2 Spanish learners requires their outputs to be complemented by a detailed phonetic analysis.

Cite

CITATION STYLE

APA

Sarymsakova, A., & Martín Rodilla, P. (2025). Assessment of L2 Spanish pronunciation accuracy via Automatic Speech Recognition. PHONICA, 21, 1–23. https://doi.org/10.1344/phonica2025.21.2

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free