Abstract
The key towards a low complexity model for convolution neural network is in controlling the number of parameters of the network and ensuring that the input representation is not extremely large. Hence, to tackle low complexity for acoustic scene classification (ASC), this paper proposed an enhanced wavelet scattering representation with a combination of mobile network modules and shuffling modules. While wavelet scattering comprises wavelet transform with multiple wavelet scales, the averaging operation to make the wavelet scattering invariant to translation limit the maximum timescale. Hence, wavelet scattering is affected by Heisenberg's Uncertainty Principle. However, creating an input representation with multiple timescales does not meet the brief of low complexity modelling. Hence, we proposed a simple mixing of the first and second order with different timescales. The result is an input representation with nearly the same dimension as the usual wavelet scattering but with enhanced multiscale. To further leverage the 'interleaved' wavelet scattering, this paper presents sub-spectral shuffling inspired by shuffling modules that use stochasticity to improve the model's generalization. Unlike channel shuffle that shuffles channel-wise and spatial shuffle that shuffles pixel-wise, sub-spectral shuffle aims at shuffling the feature maps frequency-wise with the concept of binning. Each bin is shuffled, so the high-frequency spectrum is shuffled to low-frequency spectrum position. As such, the model learns the general acoustic profile of a scene rather than memorizing what is happening at the low-frequency or high-frequency spectrum is erratic for ASC. In addition, this paper also studied temporal shuffling, which shuffles the feature maps temporal-wise, and evaluated sub-spectral shuffling, temporal shuffling, and channel shuffling individually. Our results demonstrated the superiority of sub-spectral shuffling and the modularity of shuffling modules. We then evaluate various combinations of the three shuffling modules on three acoustic scene classification datasets. Our best model combines the three shuffling modules and achieves 70.6% classification accuracy on DCASE 2021 Task 1a dataset, 82.15% on ESC-50 dataset, 81% on Urbansound8K, with 65K parameters and a size of 126.6KB. In addition, the inclusion of shuffling modules has increased the model performance. Sub-spectral shuffling is especially useful in improving logloss, a metric used to determine the confidence level of the model.
Author supplied keywords
Cite
CITATION STYLE
Kek, X. Y., Chin, C. S., & Li, Y. (2022). An Intelligent Low-Complexity Computing Interleaving Wavelet Scattering Based Mobile Shuffling Network for Acoustic Scene Classification. IEEE Access, 10, 82185–82201. https://doi.org/10.1109/ACCESS.2022.3196338
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.