VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chaofan, Hao, Peng, Cao, Xiaoge, Hao, Xiaoshuai, Cui, Shaowei, Wang, Shuo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TLA: Tactile-Language-Action Model for Contact-Rich Manipulation
by: Hao, Peng, et al.
Published: (2025)
by: Hao, Peng, et al.
Published: (2025)
OmniVTLA: Vision-Tactile-Language-Action Model with Semantic-Aligned Tactile Sensing
by: Cheng, Zhengxue, et al.
Published: (2025)
by: Cheng, Zhengxue, et al.
Published: (2025)
CLTP: Contrastive Language-Tactile Pre-training for 3D Contact Geometry Understanding
by: Ma, Wenxuan, et al.
Published: (2025)
by: Ma, Wenxuan, et al.
Published: (2025)
FG-CLTP: Fine-Grained Contrastive Language Tactile Pretraining for Robotic Manipulation
by: Ma, Wenxuan, et al.
Published: (2026)
by: Ma, Wenxuan, et al.
Published: (2026)
SpikingTac: A Miniaturized Neuromorphic Visuotactile Sensor for High-Precision Dynamic Tactile Imprint Tracking
by: Jiang, Tianyu, et al.
Published: (2026)
by: Jiang, Tianyu, et al.
Published: (2026)
DexTac: Learning Contact-aware Visuotactile Policies via Hand-by-hand Teaching
by: Zhang, Xingyu, et al.
Published: (2026)
by: Zhang, Xingyu, et al.
Published: (2026)
Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
by: Huang, Jialei, et al.
Published: (2025)
by: Huang, Jialei, et al.
Published: (2025)
Learning Bimanual Cloth Manipulation with Vision-based Tactile Sensing via Single Robotic Arm
by: Lee, Dongmyoung, et al.
Published: (2026)
by: Lee, Dongmyoung, et al.
Published: (2026)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
by: Zhang, Kaidi, et al.
Published: (2026)
by: Zhang, Kaidi, et al.
Published: (2026)
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
by: Li, Xiaoqi, et al.
Published: (2026)
by: Li, Xiaoqi, et al.
Published: (2026)
A Low-Cost Vision-Based Tactile Gripper with Pretraining Learning for Contact-Rich Manipulation
by: Liu, Yaohua, et al.
Published: (2026)
by: Liu, Yaohua, et al.
Published: (2026)
TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation
by: Huang, Yuzhe, et al.
Published: (2026)
by: Huang, Yuzhe, et al.
Published: (2026)
VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
by: Bi, Jianxin, et al.
Published: (2025)
by: Bi, Jianxin, et al.
Published: (2025)
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
by: Tu, Ruisen, et al.
Published: (2026)
by: Tu, Ruisen, et al.
Published: (2026)
Learning Tactile Insertion in the Real World
by: Palenicek, Daniel, et al.
Published: (2024)
by: Palenicek, Daniel, et al.
Published: (2024)
Survey of Vision-Language-Action Models for Embodied Manipulation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing
by: Gubernatorov, Konstantin, et al.
Published: (2026)
by: Gubernatorov, Konstantin, et al.
Published: (2026)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
by: Liang, Yuanchang, et al.
Published: (2026)
by: Liang, Yuanchang, et al.
Published: (2026)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
by: Sun, Jianli, et al.
Published: (2026)
by: Sun, Jianli, et al.
Published: (2026)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
by: Sun, Lin, et al.
Published: (2025)
by: Sun, Lin, et al.
Published: (2025)
ManiSkill-ViTac 2025: Challenge on Manipulation Skill Learning With Vision and Tactile Sensing
by: Li, Chuanyu, et al.
Published: (2024)
by: Li, Chuanyu, et al.
Published: (2024)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
by: Liu, Jinkun, et al.
Published: (2026)
by: Liu, Jinkun, et al.
Published: (2026)
Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
by: Wei, Xiangyi, et al.
Published: (2025)
by: Wei, Xiangyi, et al.
Published: (2025)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
by: Peng, Xiongfeng, et al.
Published: (2026)
by: Peng, Xiongfeng, et al.
Published: (2026)
VLP: Vision-Language Preference Learning for Embodied Manipulation
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Embodiment Transfer Learning for Vision-Language-Action Models
by: Li, Chengmeng, et al.
Published: (2025)
by: Li, Chengmeng, et al.
Published: (2025)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
Tactile Modality Fusion for Vision-Language-Action Models
by: Morissette, Charlotte, et al.
Published: (2026)
by: Morissette, Charlotte, et al.
Published: (2026)
CoFreeVLA: Collision-Free Dual-Arm Manipulation via Vision-Language-Action Model and Risk Estimation
by: Zhai, Xuanran, et al.
Published: (2026)
by: Zhai, Xuanran, et al.
Published: (2026)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
by: Li, Pengteng, et al.
Published: (2026)
by: Li, Pengteng, et al.
Published: (2026)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
A Vision-Language-Action Model for Adaptive Ultrasound-Guided Needle Insertion and Needle Tracking
by: Zhang, Yuelin, et al.
Published: (2026)
by: Zhang, Yuelin, et al.
Published: (2026)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation
by: Zhu, Juan, et al.
Published: (2026)
by: Zhu, Juan, et al.
Published: (2026)
Extremum Seeking Controlled Wiggling for Tactile Insertion
by: Burner, Levi, et al.
Published: (2024)
by: Burner, Levi, et al.
Published: (2024)
What Foundation Models can Bring for Robot Learning in Manipulation : A Survey
by: Li, Dingzhe, et al.
Published: (2024)
by: Li, Dingzhe, et al.
Published: (2024)
ETA-VLA: Efficient Token Adaptation via Temporal Fusion and Intra-LLM Sparsification for Vision-Language-Action Models
by: Wang, Yiru, et al.
Published: (2026)
by: Wang, Yiru, et al.
Published: (2026)
HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models
by: Feng, Qiuxuan, et al.
Published: (2026)
by: Feng, Qiuxuan, et al.
Published: (2026)
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
by: Xu, Haiweng, et al.
Published: (2026)
by: Xu, Haiweng, et al.
Published: (2026)
Similar Items
-
TLA: Tactile-Language-Action Model for Contact-Rich Manipulation
by: Hao, Peng, et al.
Published: (2025) -
OmniVTLA: Vision-Tactile-Language-Action Model with Semantic-Aligned Tactile Sensing
by: Cheng, Zhengxue, et al.
Published: (2025) -
CLTP: Contrastive Language-Tactile Pre-training for 3D Contact Geometry Understanding
by: Ma, Wenxuan, et al.
Published: (2025) -
FG-CLTP: Fine-Grained Contrastive Language Tactile Pretraining for Robotic Manipulation
by: Ma, Wenxuan, et al.
Published: (2026) -
SpikingTac: A Miniaturized Neuromorphic Visuotactile Sensor for High-Precision Dynamic Tactile Imprint Tracking
by: Jiang, Tianyu, et al.
Published: (2026)