Abstract
Among various activation functions, the Rectified Linear Unit (ReLU) has become the most widely adopted due to its computational simplicity and effectiveness in mitigating the vanishing-gradient problem. In this work, we investigate the advantages of employing ReLU as the activation function and establish its theoretical significance. Our analysis demonstrates that ReLU-based neural networks possess the universal approximation property. In addition, we provide a theoretical explanation for the phenomenon of neuron death in ReLU-based neural networks. We further validate the effectiveness of this explanation through empirical experiments.
Author supplied keywords
Cite
CITATION STYLE
Luo, G., Wang, X., Zhao, W., Tao, S., & Tang, Z. (2026). ReLU Neural Networks and Their Training. Mathematics, 14(1). https://doi.org/10.3390/math14010039
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.