VLH: Vision-Language-Haptics Foundation Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fuentes, Luis Francisco Moreno, Khan, Muhammad Haris, Cabrera, Miguel Altamirano, Serpiva, Valerii, Iarchuk, Dmitri, Mahmoud, Yara, Tokmurziyev, Issatay, Tsetserukou, Dzmitry
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912516380033024
author Fuentes, Luis Francisco Moreno
Khan, Muhammad Haris
Cabrera, Miguel Altamirano
Serpiva, Valerii
Iarchuk, Dmitri
Mahmoud, Yara
Tokmurziyev, Issatay
Tsetserukou, Dzmitry
author_facet Fuentes, Luis Francisco Moreno
Khan, Muhammad Haris
Cabrera, Miguel Altamirano
Serpiva, Valerii
Iarchuk, Dmitri
Mahmoud, Yara
Tokmurziyev, Issatay
Tsetserukou, Dzmitry
contents We present VLH, a novel Visual-Language-Haptic Foundation Model that unifies perception, language, and tactile feedback in aerial robotics and virtual reality. Unlike prior work that treats haptics as a secondary, reactive channel, VLH synthesizes mid-air force and vibration cues as a direct consequence of contextual visual understanding and natural language commands. Our platform comprises an 8-inch quadcopter equipped with dual inverse five-bar linkage arrays for localized haptic actuation, an egocentric VR camera, and an exocentric top-down view. Visual inputs and language instructions are processed by a fine-tuned OpenVLA backbone - adapted via LoRA on a bespoke dataset of 450 multimodal scenarios - to output a 7-dimensional action vector (Vx, Vy, Vz, Hx, Hy, Hz, Hv). INT8 quantization and a high-performance server ensure real-time operation at 4-5 Hz. In human-robot interaction experiments (90 flights), VLH achieved a 56.7% success rate for target acquisition (mean reach time 21.3 s, pose error 0.24 m) and 100% accuracy in texture discrimination. Generalization tests yielded 70.0% (visual), 54.4% (motion), 40.0% (physical), and 35.0% (semantic) performance on novel tasks. These results demonstrate VLH's ability to co-evolve haptic feedback with perceptual reasoning and intent, advancing expressive, immersive human-robot interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLH: Vision-Language-Haptics Foundation Model
Fuentes, Luis Francisco Moreno
Khan, Muhammad Haris
Cabrera, Miguel Altamirano
Serpiva, Valerii
Iarchuk, Dmitri
Mahmoud, Yara
Tokmurziyev, Issatay
Tsetserukou, Dzmitry
Robotics
We present VLH, a novel Visual-Language-Haptic Foundation Model that unifies perception, language, and tactile feedback in aerial robotics and virtual reality. Unlike prior work that treats haptics as a secondary, reactive channel, VLH synthesizes mid-air force and vibration cues as a direct consequence of contextual visual understanding and natural language commands. Our platform comprises an 8-inch quadcopter equipped with dual inverse five-bar linkage arrays for localized haptic actuation, an egocentric VR camera, and an exocentric top-down view. Visual inputs and language instructions are processed by a fine-tuned OpenVLA backbone - adapted via LoRA on a bespoke dataset of 450 multimodal scenarios - to output a 7-dimensional action vector (Vx, Vy, Vz, Hx, Hy, Hz, Hv). INT8 quantization and a high-performance server ensure real-time operation at 4-5 Hz. In human-robot interaction experiments (90 flights), VLH achieved a 56.7% success rate for target acquisition (mean reach time 21.3 s, pose error 0.24 m) and 100% accuracy in texture discrimination. Generalization tests yielded 70.0% (visual), 54.4% (motion), 40.0% (physical), and 35.0% (semantic) performance on novel tasks. These results demonstrate VLH's ability to co-evolve haptic feedback with perceptual reasoning and intent, advancing expressive, immersive human-robot interactions.
title VLH: Vision-Language-Haptics Foundation Model
topic Robotics
url https://arxiv.org/abs/2508.01361