SPRING: Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph

N/ACitations
Citations of this article
10Readers
Mendeley users who have this article in their library.

Abstract

Existing multimodal conversation agents have shown impressive abilities to locate absolute positions or retrieve attributes in simple scenarios, but they fail to perform well when complex relative positions and information alignments are involved, which poses a bottleneck in response quality. In this paper, we propose a Situated Conversation Agent PRetrained with Multimodal Questions from INcremental Layout Graph (SPRING) with abilities of reasoning multi-hops spatial relations and connecting them with visual attributes in crowded situated scenarios. Specifically, we design two types of Multimodal Question Answering (MQA) tasks to pretrain the agent. All QA pairs utilized during pretraining are generated from novel Incremental Layout Graphs (ILG). QA pair difficulty labels automatically annotated by ILG are used to promote MQA-based Curriculum Learning. Experimental results verify the SPRING's effectiveness, showing that it significantly outperforms state-of-the-art approaches on both SIMMC 1.0 and SIMMC 2.0 datasets. We release our code and data at Github LYX0501/SPRING repository.

Cite

CITATION STYLE

APA

Long, Y., Hui, B., Ye, F., Li, Y., Han, Z., Yuan, C., … Wang, X. (2023). SPRING: Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, AAAI 2023 (Vol. 37, pp. 13309–13317). AAAI Press. https://doi.org/10.1609/aaai.v37i11.26562

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free