SSF4VSU: A Self-Supervised Synergetic Framework for Visual Scene Understanding

1Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Visual scene understanding involves multiple tasks, including single-object tracking (SOT), multi-object tracking (MOT), video object segmentation (VOS), and its Multi-Object Tracking and Segmentation (MOTS) variant, which have traditionally been studied in isolation. Such fragmentation leads to task-specific architectures that must be retrained for each new scenario and rely heavily on annotated data. This work proposes a synergetic model called SSF4VSU that simultaneously addresses SOT, MOT, VOS, and MOTS. SSF4VSU employs a shared backbone with a unified embedding space for different tasks, a Temporal Attention Module (TAM) to align features across frames and resist occlusions, a Temporal Consistency Module (TCM) to enforce smooth identity preservation, and a self-supervised learning (SSL) branch that leverages unlabeled video sequences for improved generalization. Training follows a multi-task curriculum with dynamic loss balancing, demonstrating efficiency and scalability. A comprehensive evaluation across six public benchmarks shows that SSF4VSU surpasses specialised and unified state-of-the-art models. On SOT benchmarks it achieves 74.7% success AUC and 80.4% precision on LaSOT, and 85.8% AUC and 84.3% precision on TrackingNet, matching or exceeding recent specialised trackers. For MOT, it records 82.9% MOTA and 83.3% IDF1 on the MOT17 dataset, while on BDD100K it generalises to driving scenes without task-specific tuning. On VOS benchmarks, SSF4VSU delivers 93.3% J\& F on DAVIS-2016 and 89.0% J & F on DAVIS-2017, surpassing strong segmentation methods and unified baselines. In the MOTS setting it achieves 69.0% sMOTSA on MOTS20 and 31.2% mMOTSA on BDD100K MOTS. Ablation studies reveal that TAM is critical for accurate localisation, TCM is indispensable for identity stability, and SSL consistently improves generalization. The results demonstrate that a single unified model can perform diverse tracking and segmentation tasks without sacrificing accuracy or efficiency. By reducing reliance on labelled data, improving temporal reasoning, and ensuring interpretability, SSF4VSU advances the state of visual scene understanding and lays the foundation for general-purpose, multi-task video analysis.

Cite

CITATION STYLE

APA

Hassan, S., Mujtaba, G., Ullah, H., Imran, A. S., Soylu, A., & Ullah, M. (2025). SSF4VSU: A Self-Supervised Synergetic Framework for Visual Scene Understanding. IEEE Access, 13, 197544–197561. https://doi.org/10.1109/ACCESS.2025.3634778

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free