10 Open Challenges Steering the Future of Vision-Language-Action Models

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Due to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly prevalent in the embodied AI arena, following the widespread success of their precursors—LLMs and VLMs. In this paper, we discuss 10 principal milestones in the ongoing development of VLA models—multimodality, reasoning, data, evaluation, cross-robot action generalization, efficiency, whole-body coordination, safety, agents, and coordination with humans. Furthermore, we discuss the emerging trends of using spatial understanding, modeling world dynamics, post training, and data synthesis—all aiming to reach these milestones. Through these discussions, we hope to bring attention to the research avenues that may accelerate the development of VLA models into wider acceptability.

Cite

CITATION STYLE

APA

Poria, S., Majumder, N., Hung, C. Y., Bagherzadeh, A. A., Li, C., Kwok, K., … Hsu, D. (2026). 10 Open Challenges Steering the Future of Vision-Language-Action Models. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 40, pp. 39771–39779). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/aaai.v40i46.41333

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free