Abstract
Tactile feedback is generally recognized to be crucial for effective interaction with the physical world. However, state-of-the-art Vision-Language-Action (VLA) models often lack the ability to interpret and use tactile signals, limiting their effectiveness in contact-rich tasks. Incorporating tactile feedback into these systems is challenging due to the absence of large multi-modal datasets. We present VLA-Touch, an approach that enhances generalist robot policies with tactile sensing without fine-tuning the base VLA with tactile data. Our method introduces two key innovations: (1) a pipeline that leverages a pretrained tactile-language model to provide semantic tactile feedback for high-level task planning, and (2) a diffusion-based controller that refines VLA-generated actions with tactile signals for contact-rich manipulation. Through real-world experiments, we demonstrate that our dual-level integration of tactile feedback improves task planning success rate while enhancing execution precision.
Author supplied keywords
Cite
CITATION STYLE
Bi, J., Ma, K. Y., Hao, C., Zheng, M. S., & Soh, H. (2026). VLA-Touch: Enhancing Vision-Language-Action Model With Dual-Level Tactile Feedback. IEEE Robotics and Automation Letters, 11(7), 8487–8494. https://doi.org/10.1109/LRA.2026.3692345
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.