TEMPORAL FUSION STRATEGY FOR VIOLENCE DETECTION: UTILISING CONVOLUTIONAL AND LSTM NEURAL NETWORKS FOR SURVEILLANCE VIDEOS

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

In the latest intelligent cities, there is a pursuit for the utmost degree of automation and integration of services. One of the major challenges in the surveillance industry is the need to automate real-time video analysis to identify critical cases. This paper introduces sophisticated models using Convolutional Neural Networks (CNN), specifically MobileNet V3, VGG16, and InceptionV3 networks, as well as networks using LSTM and feedforward networks. These models are designed to accurately categorise videos into two completely separate classes, namely: (“Non-Violence” and “Violence”). The RLVS database is used for this classification task. Various data representations are used by Temporal Fusion approaches. The highest attained outcome was an Accuracy of 91.03 %, and an F1-score of 90.90 %, which is superior to the results obtained in similar research performed on the same database for achieving the goal of recognising actions that are violent in Surveillance Videos.

Cite

CITATION STYLE

APA

Merit, K., Beladgham, M., & Taleb-Ahmed, A. (2025). TEMPORAL FUSION STRATEGY FOR VIOLENCE DETECTION: UTILISING CONVOLUTIONAL AND LSTM NEURAL NETWORKS FOR SURVEILLANCE VIDEOS. Acta Polytechnica, 65(3), 306–319. https://doi.org/10.14311/AP.2025.65.0306

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free