Accent-robust speech recognition for English in low-resource settings using Manifold Mixup

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

We adapt Manifold Mixup theory for accent-robust end-to-end (E2E) Automatic speech recognition (ASR). Accent-variation between a source and target constitutes a domain-mismatch scenario. Manifold Mixup allows cross-domain robustness where a model trained on a source accent generalizes to target accents. We propose a 2-stage training mechanism with manifold mixup using one source accent. Stage 1 is a mixup-enabled cross-entropy based framewise character recognition model. Stage 2 is a Connectionist Temporal Classification (CTC)-loss based E2E ASR model using Stage 1 weights. We show that this model generalizes to unseen accents without any fine-tuning. This is studied for accented English from Indic-TIMIT corpus (6 Indic accents) and Common Voice corpus accent groups UKI (England, Ireland), Oriental (India, Malaysia), NorthAM (USA, Canada), African and ANZ (Australia, New Zealand). This is also studied with another Indian English corpus Svarah, the American English TIMIT corpus and the open audiobook English corpus of Librispeech. The proposed framework, using a Hindi-mixup model, offers absolute gains of around 2% over a non-mixup baseline on unseen test accents.

Cite

CITATION STYLE

APA

Banerjee, T., & Ramasubramanian, V. (2025). Accent-robust speech recognition for English in low-resource settings using Manifold Mixup. Eurasip Journal on Audio, Speech, and Music Processing, 2025(1). https://doi.org/10.1186/s13636-025-00435-0

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free