A Methodology for Evaluating Multimodal Referring Expression Generation for Embodied Virtual Agents

4Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Robust use of definite descriptions in a situated space often involves recourse to both verbal and non-verbal modalities. For IVAs, virtual agents designed to interact with humans, the ability to both recognize and generate non-verbal and verbal behavior is a critical capability. To assess how well an IVA is able to deploy multimodal behaviors, including language, gesture, and facial expressions, we propose a methodology to evaluate the agent's capacity to generate object references in a situational context, using the domain of multimodal referring expressions as a use case. Our contributions include: 1) developing an embodied platform to collect human referring expressions while communicating with the IVA. 2) comparing human and machine-generated references in terms of evaluable properties using subjective and objective metrics. 3) reporting preliminary results from trials that aimed to check whether the agent can retrieve and disambiguate the object the human referred to, if the human has the ability to correct misunderstanding using language, deictic gesture, or both; and human ease of use while interacting with the agent.

Cite

CITATION STYLE

APA

Alalyani, N., & Krishnaswamy, N. (2023). A Methodology for Evaluating Multimodal Referring Expression Generation for Embodied Virtual Agents. In ACM International Conference Proceeding Series (pp. 164–173). Association for Computing Machinery. https://doi.org/10.1145/3610661.3616548

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free