Abstract
The paper describes a system for automatic evaluation of speech quality based on statistical analysis of differences in spectral properties, prosodic parameters, and time structuring within the speech signal. The proposed system was successfully tested in evaluation of sentences originating from male and female voices and produced by a speech synthesizer using the unit selection method with two different approaches to prosody manipulation. The experiments show necessity of all three types of speech features for obtaining correct, sharp, and stable results. A detailed analysis shows great influence of the number of statistical parameters on correctness and precision of the evaluated results. Larger size of the processed speech material has a positive impact on stability of the evaluation process. Final comparison documents basic correlation with the results obtained by the standard listening test.
Author supplied keywords
Cite
CITATION STYLE
Přibil, J., Přibilová, A., & Matoušek, J. (2018). Automatic evaluation of synthetic speech quality by a system based on statistical analysis. In Lecture Notes in Computer Science (Vol. 11107 LNAI, pp. 315–323). Springer Verlag. https://doi.org/10.1007/978-3-030-00794-2_34
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.