Abstract
Objective Pedestrian trajectory prediction is essential for such domains like unmanned vehicles, security surveillance, and social robotics nowadays. Trajectory prediction is beneficial for computer systems to perform better decision making and planning to some extent. Current methods are focused on pedestrian trajectory information, and scene elements-related spatial constraints on pedestrian motion in the same space are challenged to explain human-to-human social interactions further,in which future location of pedestrians cannot be located in building walls, and pedestrians at building corners undergo large velocity direction deflections due to cornering behavior. The pathways can be focused on the integrated scene information, for which the scene image is melted into a one-dimensional vector and merged with the trajectory information. Two-dimensional spatial signal of the scene will be distorted and it cannot be intuitively explained according to the modulating effect of the scene on pedestrian motion. To build a spatiotemporal graph representation of pedestrians, recent graph neural network (G N N) is used to develop a method based on graph attention network(G A T),in which pedestrians are as the graph nodes, trajectory features as the node attributes,and pedestrians-between spatial interactions are as the edges in the graph. These sorts of methods can be used to focus on pedestrians-between social interactions in the global scale. However, for crowded scenes, graph attention mechanism may not be able to assign appropriate weights to each pedestrian accurately, resulting in poor algorithm accuracy. To resolve the two problems mentioned above, we develop a scene constraints-based spatiotemporal graph convolutional network,called Scene-STGCNN,which aggregates pedestrian motion status with a graph convolutional neural network for local interactions, and it achieves accurate aggregation of pedestrian motion status with a small number of parameters. At the same time, we design a scene-based fine-tuning module to explicitly model the modulating effect of scenes on pedestrian motion with the information of neighboring scene changes as inpul. Method Scene-STGCNN consists of a motion module, a scene-based fine-tuning module, spatiotemporal convolution, and spatiotemporal extrapolation convolution. For the motion module, the graph convolution is a 1 x 1 coresized convolutional neural network (C N N) layer for embedding pedestrian velocity information. The residual convolution is composed of C N N layer of 1 X 1 kernel size and BatchNorm (BN) layer. Temporal convolution is organized of B N layer, P R e L U layer, 3 X 1 core-sized C N N layer, B N layer and Dropout layer as well. The motion module takes the pedestrian velocity spatiotemporal graph and the scene mask matrix as input, in which CNN-based pedestrian velocity spatiotemporal graph is encoded and the pedestrian spatiotemporal features of existing multiple frames are fused. For the scene-based finetuning module,temporal neighboring scene change information is first introduced to generate the scene-based pedestrian spatiotemporal m a p,and the embedding of the pedestrian spatiotemporal map by scene convolution is then performed to obtain the scene mask matrix,which is used to make Hadamard products with the intermediate motion features in the motion module. The real-time regulation role of the scene on pedestrians can be explicitly modeling further. Spatiotemporal convolution as a transition coding network consists of two temporal gating units and a spatial convolution,which is used to enhance the temporal correlation and contextual spatial dependence of pedestrian motion. A two-dimensional Gaussian distribution-related trajectory distribution is generated in terms of temporal extrapolation convolution. The kernel density estimation-based negative log-likelihood as the loss function will enhance the multimodality of the Scene-STGCNN prediction distribution while the prediction loss is optimized. Result Experiments are carried out to compare with the other related seven popular methods on the publicly available datasets E T H (including E T H and H O T E L) and U C Y (including UNIV, Z A R A 1, and ZARA2). The average displacement error (ADE) values are optimized by 1 2 %, and the final displacement error (FDE) values are optimized by 9 % in terms of average values. Ablation experiments are used to verify the effectiveness of the scene-based fine-tuning module, and the results demonstrate that the scene-based fine-tuning module can effectively model the modulation effect of the scene on pedestrian trajectory, and the prediction error of the algorithm is optimized as well. In addition, qualitative analysis is focused on the issues of Scene-STGCNN-captured inherent patterns of pedestrian motion and the involved prediction distribution. The visualization results show that Scene-STGCNN can be used to learn the pedestrian motion patterns effectively while maintaining accurate predictions. Conclusion we facilitate a pedestrian trajectory prediction model, called Scene-STGCNN, which can fuse scene information with trajectory features effectively through a scene-based fine-tuning module. Furthermore, Scene-STGCNN potentials can be focused on scene information-related pedestrian trajectory prediction method to a certain extent via modeling the modulation effect of scene on pedestrian motion.
Author supplied keywords
Cite
CITATION STYLE
Chen, H., & Ji, Q. (2023). Scene-constrained spatial-temporal graph convolutional network for pedestrian trajectory prediction. Journal of Image and Graphics, 22(8), 3163–3175. https://doi.org/10.11834/jig.221027
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.