Abstract
3D Convolution Neural Networks (CNNs), an important deep learning model, has good performance in recognizing actions in videos. When recognizing actions from videos, 3D CNNs usually down-sample in temporal dimension, leading to loss of the temporal information. To obtain more temporal information from the videos, this work proposed a new model based on the Inflated 3D ConvNet (I3D), named as I3D-T. Instead of using down-sample in temporal dimension, the proposed model applied the dilated convolution in temporal dimension to enlarge the receptive field. At the same time, a non-local feature gating block was designed in the model to learn the correlations between different feature maps. The experimental results showed that the proposed I3D-T has the state-of-art performance. Using RGB frames as input, the action recognition accuracies are respectively 95% and 74.8% in public dataset of UCF101 and HMDB-51.
Author supplied keywords
Cite
CITATION STYLE
Xu, Y., Feng, Y., Xie, Z., Xie, M., & Luo, W. (2020). Action recognition using high temporal resolution 3D neural network based on dilated convolution. IEEE Access, 8, 165365–165372. https://doi.org/10.1109/ACCESS.2020.3022407
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.