Applicability of End-to-End Deep Neural Architecture to Sinhala Speech Recognition

  • Gamage B
  • Pushpananda R
  • Nadungodage T
  • et al.
N/ACitations
Citations of this article
10Readers
Mendeley users who have this article in their library.

Abstract

This research presents a study on the application of end-to-end deep learning models for Automatic Speech Recognition in the Sinhala language, which is characterized by its high inflection and limited resources.We explore two e2e architectures, namely the e2e Lattice-Free Maximum Mutual Information model and the Recurrent Neural Network model, using a restricted dataset. Statistical models with 40 hours of training data are established as baselines for evaluation. Our pretrained endto-end Automatic Speech Recognition models achieved a Word Error Rate of 23.38% by far the best word-error-rate achieved for low resourced Sinhala Language. Our models demonstrate greater contextual independence and faster processing, making them more suitable for general-purpose speech-to-text translation in Sinhala.

Cite

CITATION STYLE

APA

Gamage, B., Pushpananda, R., Nadungodage, T., & Weerasinghe, R. (2024). Applicability of End-to-End Deep Neural Architecture to Sinhala Speech Recognition. International Journal on Advances in ICT for Emerging Regions (ICTer), 17(1), 17–21. https://doi.org/10.4038/icter.v17i1.7273

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free