KnowZRel: Common Sense Knowledge-Based Zero-Shot Relationship Retrieval for Generalized Scene Graph Generation

9Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

A scene graph is a key image representation in visual reasoning. The generalizability of scene graph generation (SGG) methods is crucial for reliable reasoning and real-world applicability. However, imbalanced training datasets limit this, underrepresenting meaningful visual relationships. Current SGG methods using external knowledge sources face limitations due to these imbalances or restricted relationship coverage, impacting their reasoning and generalization capabilities. We propose a novel neurosymbolic approach that integrates data-driven object detection with heterogeneous knowledge graph-based object refinement and zero-shot relationship retrieval, highlighting the loosely coupled synergy between neural and symbolic components. This combination addresses the limitations of imbalanced training datasets in SGG and enables effective prediction of unseen visual relationships. Objects are detected using a regionbased deep neural network and refined based on their positional and structural similarity, followed by retrieval of pairwise visual relationships using a heterogeneous knowledge graph. The redundant and irrelevant visual relationships are discarded based on the similarity of relationship labels and node embeddings. Finally, the visual relationships are interlinked to generate the scene graph. The employed heterogeneous knowledge graph combines diverse knowledge sources, offering rich common sense knowledge about objects and their interactions in the world. Our method, evaluated using the benchmark visual genome (VG) dataset and zero-shot recall (zR@K) metric, shows a 59.96% improvement over existing state-of-the-art methods, highlighting its effectiveness in generalized SGG. The object refinement step effectively improved the object detection performance by 57.1%. Additional evaluation using the GQA dataset confirms the cross-dataset generalizability of our method. We also compared various knowledge sources and embedding models to determine an optimal combination for zero-shot SGG.

Cite

CITATION STYLE

APA

Khan, M. J., Breslin, J. G., & Curry, E. (2025). KnowZRel: Common Sense Knowledge-Based Zero-Shot Relationship Retrieval for Generalized Scene Graph Generation. IEEE Transactions on Artificial Intelligence, 6(12), 3184–3194. https://doi.org/10.1109/TAI.2025.3544177

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free