Significance of GMM-UBM based Modelling for Indian Language Identification

V. Ravi Kumar; Hari Krishna Vydana; Anil Kumar Vuppala

Conference ProceedingsOPEN ACCESS

Significance of GMM-UBM based Modelling for Indian Language Identification

Procedia Computer Science (2015) 54 231-236

DOI: 10.1016/j.procs.2015.06.027

15Citations

23Readers

Abstract

Most of the Indian languages are originated from Devanagari, the script of the Sanskrit language. In-spite of similarity in phoneme sets, every language its own influence on the phonotactic constraints of speech in that language. A modelling technique that is capable of capturing the slightest variations imparted by the language is a pre-requisite for developing a language identification system (LID). Use of Gaussian mixture modelling technique with a large number of mixture components demands a large training data for each language class, which is hard to collect and handle. In this work, phonotactic variations imparted by the different languages are modelled using Gaussian mixture modelling with a universal background model (GMM-UBM) technique. In GMM-UBM based modelling certain amount of data from all the language classes is pooled to develop a universal background model (UBM) and the model is adapted to each class. Spectral features (MFCC) are employed to represent the language specific phonotactic information of speech in different languages. During the present study, LID systems are developed using the speech samples from IITKGP-MLILSC. In this work, performance of the proposed GMM-UBM based LID system is compared with conventional GMM based LID system. An average improvement of 7-8% is observed due to the use of UBM-based modelling of developing a LID system.

Author supplied keywords

Cite

CITATION STYLE

APA

Kumar, V. R., Vydana, H. K., & Vuppala, A. K. (2015). Significance of GMM-UBM based Modelling for Indian Language Identification. In Procedia Computer Science (Vol. 54, pp. 231–236). Elsevier. https://doi.org/10.1016/j.procs.2015.06.027

Significance of GMM-UBM based Modelling for Indian Language Identification

Abstract

Author supplied keywords

Cite

Register to see more suggestions