Abstract
Speech separation, as an important research direction in audio signal processing, has been widely studied by the academic community since its emergence in the mid-1990s. In recent years, with the rapid development of deep neural network technology, speech processing based on deep neural networks has shown outstanding performance in speech separation. While existing studies have surveyed the application of deep neural networks in speech separation from multiple dimensions including learning paradigms, model architectures, loss functions, and training strategies, current achievements still lack systematic comprehension of the field’s developmental trajectory. To address this, this paper focuses on single-channel supervised speech separation tasks, proposing a technological evolution path “U-Net–TasNet–Transformer–Mamba” as the main thread to systematically analyze the impact mechanisms of core architectural designs on separation performance across different stages. By reviewing the transition process from traditional methods to deep learning paradigms and delving into the improvements and integration of deep learning architectures at various stages, this paper summarizes milestone achievements, mainstream evaluation frameworks, and typical datasets in the field, while also providing prospects for future research directions. Through this detailed-focused review perspective, we aim to provide researchers in the speech separation field with a clearly articulated technical evolution map and practical reference.
Author supplied keywords
Cite
CITATION STYLE
Wang, Z., & Luo, Z. (2025, November 1). Speech Separation Using Advanced Deep Neural Network Methods: A Recent Survey. Big Data and Cognitive Computing. Multidisciplinary Digital Publishing Institute (MDPI). https://doi.org/10.3390/bdcc9110289
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.