Comparative Analysis of Automatic Speech Recognition Fine-Tuning Strategies for Speech from Cochlear Implant Users

1Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Although automatic speech recognition technology has become widespread, it still exhibits limited performance when processing speech from cochlear implant (CI) users. This limitation serves as a barrier that hinders CI users' accessibility to digital technology. To address this issue, a comparative study of fine-tuning strategies was conducted to effectively adapt Whisper, a general-purpose speech recognition model, to CI users' speech. Specifically, the performance of full fine-tuning, selective fine-tuning, adapter, and LoRA were evaluated based on Korean CI user's speech dataset. The experimental results showed that all the fine-tuning approaches improved recognition performance compared to the baseline Whisper model. Notably, LoRA-encoder approach, which involved training only 2.15% of the total parameters, achieved the best performance with a character error rate of 11.57%, demonstrating superior performance and efficiency. Furthermore, strategies that fine-tuned only the encoder consistently showed higher performance than those that adjusted the decoder, confirming that the encoder's role is crucial in modeling the unique acoustic characteristics of CI users' speech.

Cite

CITATION STYLE

APA

Yoon, S., Kim, H., Kim, K., & Lee, S. (2026). Comparative Analysis of Automatic Speech Recognition Fine-Tuning Strategies for Speech from Cochlear Implant Users. IEEE Signal Processing Letters, 33, 236–240. https://doi.org/10.1109/LSP.2025.3640524

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free