VLA-Touch: Enhancing Vision-Language-Action Model With Dual-Level Tactile Feedback

1Citations
Citations of this article
11Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Tactile feedback is generally recognized to be crucial for effective interaction with the physical world. However, state-of-the-art Vision-Language-Action (VLA) models often lack the ability to interpret and use tactile signals, limiting their effectiveness in contact-rich tasks. Incorporating tactile feedback into these systems is challenging due to the absence of large multi-modal datasets. We present VLA-Touch, an approach that enhances generalist robot policies with tactile sensing without fine-tuning the base VLA with tactile data. Our method introduces two key innovations: (1) a pipeline that leverages a pretrained tactile-language model to provide semantic tactile feedback for high-level task planning, and (2) a diffusion-based controller that refines VLA-generated actions with tactile signals for contact-rich manipulation. Through real-world experiments, we demonstrate that our dual-level integration of tactile feedback improves task planning success rate while enhancing execution precision.

Cite

CITATION STYLE

APA

Bi, J., Ma, K. Y., Hao, C., Zheng, M. S., & Soh, H. (2026). VLA-Touch: Enhancing Vision-Language-Action Model With Dual-Level Tactile Feedback. IEEE Robotics and Automation Letters, 11(7), 8487–8494. https://doi.org/10.1109/LRA.2026.3692345

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free