Abstract
In recent years, deep-learning-based speech synthesis has garnered substantial attention, achieving remarkable advancements in generating human-like speech. However, synthesized speech often lacks naturalness, primarily because models excessively depend on fine-grained text–speech alignment. To address this issue, we propose MixDiff-TTS, a novel non-autoregressive model. MixDiff-TTS incorporates a linguistic encoder based on a mixture alignment mechanism, which combines word-level hard alignment with phoneme-level soft alignment. This design reduces reliance on fine-grained alignment, enabling the model to handle ambiguous phonetic boundaries more robustly. Additionally, we introduce a Word-to-Phoneme Attention module with a relative position bias mechanism to improve the model’s capacity for processing long text sequences. We evaluate the performance of MixDiff-TTS on the LJSpeech dataset. The experimental results show that MixDiff-TTS scores 0.507 for SSIM (Structural Similarity Index) and 6.652 for MCD (Mel Cepstral Distortion). This suggests that the synthesized speech is closer to real speech in spectral structure and exhibits lower spectral distortion than state-of-the-art baselines (such as FastSpeech2 and DiffSpeech). MixDiff-TTS also achieves a MOS (Mean Opinion Score) of 3.95, which is close to that of real speech. These results indicate that MixDiff-TTS can synthesize speech with high naturalness and quality. Ablation studies demonstrate the effectiveness of our method.
Author supplied keywords
Cite
CITATION STYLE
Long, Y., Yang, K., Ma, Y., & Yang, Y. (2025). MixDiff-TTS: Mixture Alignment and Diffusion Model for Text-to-Speech. Applied Sciences (Switzerland), 15(9). https://doi.org/10.3390/app15094810
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.