Abstract
Due to the high cost of large-scale strong labeling, sound event detection (SED) using only weakly-labeled and unlabeled data has drawn increasing attention in recent years. To exploit large amount of unlabeled in-domain data efficiently, we applied three semi-supervised learning strategies: Interpolation consistency training (ICT), shift consistency training (SCT), and weakly pseudo-labeling. In addition, we propose FP-CRNN, a convolutional recurrent neural network (CRNN) which contains feature-pyramid (FP) components, to leverage temporal information by utilizing features at different scales. Experiments were conducted on DCASE 2020 task 4. In terms of event-based F-measure, these approaches outperform the official baseline system, at 34.8%, with the highest Fmeasure of 48.0% achieved by an FP-CRNN that was trained with the combination of all three strategies.
Author supplied keywords
Cite
CITATION STYLE
Koh, C. Y., Chen, Y. S., Liu, Y. W., & Bai, M. R. (2021). Sound event detection by consistency training and pseudo-labeling with feature-pyramid convolutional recurrent neural networks. In ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings (Vol. 2021-June, pp. 376–380). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1109/ICASSP39728.2021.9414350
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.