Placing objects in gesture space: Toward incremental interpretation of multimodal spatial descriptions

Ting Han; Casey Kennington; David Schlangen

Conference ProceedingsOPEN ACCESS

Placing objects in gesture space: Toward incremental interpretation of multimodal spatial descriptions

32nd AAAI Conference on Artificial Intelligence, AAAI 2018 (2018) 5157-5164

DOI: 10.1609/aaai.v32i1.11974

1Citations

12Readers

Abstract

When describing routes not in the current environment, a common strategy is to anchor the description in configurations of salient landmarks, complementing the verbal descriptions by “placing” the non-visible landmarks in the gesture space. Understanding such multimodal descriptions and later locating the landmarks from real world is a challenging task for the hearer, who must interpret speech and gestures in parallel, fuse information from both modalities, build a mental representation of the description, and ground the knowledge to real world landmarks. In this paper, we model the hearer's task, using a multimodal spatial description corpus we collected. To reduce the variability of verbal descriptions, we simplified the setup to use simple objects as landmarks. We describe a real-time system to evaluate the separate and joint contributions of the modalities. We show that gestures not only help to improve the overall system performance, even if to a large extent they encode redundant information, but also result in earlier final correct interpretations. Being able to build and apply representations incrementally will be of use in more dialogical settings, we argue, where it can enable immediate clarification in cases of mismatch.

Cite

CITATION STYLE

APA

Han, T., Kennington, C., & Schlangen, D. (2018). Placing objects in gesture space: Toward incremental interpretation of multimodal spatial descriptions. In 32nd AAAI Conference on Artificial Intelligence, AAAI 2018 (pp. 5157–5164). AAAI press. https://doi.org/10.1609/aaai.v32i1.11974

Placing objects in gesture space: Toward incremental interpretation of multimodal spatial descriptions

Abstract

Cite

Register to see more suggestions