Towards Fair and Inclusive Speech Recognition for Stuttering: Community-led Chinese Stuttered Speech Dataset Creation and Benchmarking

9Citations
Citations of this article
11Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Despite the widespread adoption of Automatic Speech Recognition (ASR) models in voice-operated products and conversational AI agents, current ASR models perform poorly for people who stutter. One primary cause of the performance disparity is the lack of representative stuttered speech data during the development of ASR models. This work introduces the first stuttered speech dataset in Mandarin Chinese, created by a grassroots community of Chinese-speaking people who stutter to facilitate the development of inclusive and fair speech AI. Collected from 72 speakers with a wide range of stuttering characteristics, this dataset contains speech samples of both spontaneous conversations and voice command dictations from each speaker. Our analysis of the dataset shows the diversity and variability of stuttered utterances captured, highlighting its unique value in authentically representing the stuttering community in AI data. Leveraging this dataset, we benchmark popular ASR models to understand their potential biases against disfluent speech.

Cite

CITATION STYLE

APA

Li, Q., & Wu, S. (2024). Towards Fair and Inclusive Speech Recognition for Stuttering: Community-led Chinese Stuttered Speech Dataset Creation and Benchmarking. In Conference on Human Factors in Computing Systems - Proceedings. Association for Computing Machinery. https://doi.org/10.1145/3613905.3650950

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free