Abstract
This paper addresses the detection and classification of mixed-critical events, such as fire incidents, traffic accidents, and violence-related events in urban surveillance systems. It proposes a novel multimodal lightweight framework for these applications. Traditional event detection techniques often use single-modality systems, such as Convolutional Neural Networks (CNNs), Vision Transformer (ViTs), and the You Only Look Once (YOLO) variants. These techniques face difficulties when it comes to effectively grasping both the severity and contextual complexity of events. This paper addresses the aforementioned gaps with a VLM-based multimodal framework using Bootstrapped Language-Image Pretraining (BLIP) and Uform-Gen to generate descriptive captions and pretext embeddings. With this approach, the comprehension of context extends beyond a single event and captures the relationship across multiple event types. The preliminary results show that the proposed approach integrates visual and textual data more efficiently than previous methods, thus enhancing the accuracy of classification. It has also been shown that BLIP yields 98.04% on pretext embeddings and 94.08% on captions, while Uform-Gen achieves 95.30%, 94.40% on pretext embeddings and captions, respectively, on the traffic dataset, confirming the contextual modeling and comprehensive event understanding provided by our proposed multimodal learning approach. Our work also highlights AI-led, multimodal frameworks for event detection that incorporate spatial and situational reasoning, thus bridging the gap between automated real-time incident monitoring and practical deployment.
Author supplied keywords
Cite
CITATION STYLE
Sadhwani, S., Shamsi, J. A., Khan, M. B., Bawany, N. Z., & Syed, H. J. (2025). Real-Time Detection of Mixed-Critical Events Using Vision-Language Models. IEEE Access, 13, 181363–181384. https://doi.org/10.1109/ACCESS.2025.3622638
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.