Abstract
The rapid growth of urban environments and vehicle usage has made it more challenging to interpret and extract insights from surveillance videos. Surveillance traffic systems produce a huge amount of video, rendering manual screening time-consuming and inefficient. To meet this severe challenge, our work presents an efficient and scalable video summarization system tailored for the quick and precise analysis of Indian traffic surveillance video. The framework combines cutting-edge deep learning object detection models such as YOLOv8, YOLO11, Faster R-CNN, and RetinaNet, which were trained on the Indian Road User Vehicle Dataset (IRUVD) a detailed dataset reflecting varied road environments peculiar to India and tested on a custom video dataset captured in Bangalore, India. Following a thorough comparative study in terms of performance metrics such as mAP50, precision, recall, training stability, and inference speed, YOLO11 emerged as the best performing model for deployment in real-world applications due to its balance between accuracy and compute efficiency. In addition to object detection, the system uses sophisticated masking methods to filter and emphasize vehicle classes or objects of interest to support focused video summarization. Several masking approaches were created like black and white masking, blackout, Gaussian blur, telea inpainting, and selective blur to provide adaptability according to traffic enforcement, urban planning and anomaly detection analysis requirements for Indian traffic scenario. Experimental evaluation of the proposed system showed excellent object detection precision and recall scores of 0.9504, with a mAP@0.5 value of 0.9688 and an overall Fitness Score of 90.14%. Frame selection also achieved good results with an F1-score of 0.9705 and an accuracy of 0.9750, thus assuring the systems reliability in selecting relevant vehicle classes and frames for efficient video summarization. Visual masking was assessed using PSNR and MSE metrics keeping selective texture blurring masking technique as the base. It maintained contextual integrity with PSNR of 31.72 while on the other hand structural segmentation achieved higher concealment with PSNR of 5.59, indicating the framework’s flexibility across privacy and use-case driven approach.
Author supplied keywords
Cite
CITATION STYLE
Saraff, A., Gite, A., Joshi, A., Mohana, Kumar, P. R., Sreelakshmi, K., & Shankar, T. (2025). Indian Traffic Surveillance Video Summarization Using YOLO and Multi-Level Masking. IEEE Access, 13, 171371–171385. https://doi.org/10.1109/ACCESS.2025.3616267
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.