Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Zhixiang, Tian, Di, Guan, Runwei, Mu, Yanzhou, Sun, Xiaolou, Liang, Shaofeng, Liu, Daizong, Huang, Tao, Yue, Yutao, Ding, Henghui, Fang, Bin, Zhou, Alex, Han, Qing-Long, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cognitive Disentanglement for Referring Multi-Object Tracking
by: Liang, Shaofeng, et al.
Published: (2025)
by: Liang, Shaofeng, et al.
Published: (2025)
Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios
by: Chen, Chenglizhao, et al.
Published: (2025)
by: Chen, Chenglizhao, et al.
Published: (2025)
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
by: Hou, Minghui, et al.
Published: (2025)
by: Hou, Minghui, et al.
Published: (2025)
AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild
by: Sun, Xiaolou, et al.
Published: (2026)
by: Sun, Xiaolou, et al.
Published: (2026)
Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature Fusion
by: Liu, Bin, et al.
Published: (2026)
by: Liu, Bin, et al.
Published: (2026)
UniBEVFusion: Unified Radar-Vision BEVFusion for 3D Object Detection
by: Zhao, Haocheng, et al.
Published: (2024)
by: Zhao, Haocheng, et al.
Published: (2024)
ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave Radar
by: Guan, Runwei, et al.
Published: (2023)
by: Guan, Runwei, et al.
Published: (2023)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
by: Zhang, Kaidi, et al.
Published: (2026)
by: Zhang, Kaidi, et al.
Published: (2026)
Vision-Language Navigation with Embodied Intelligence: A Survey
by: Gao, Peng, et al.
Published: (2024)
by: Gao, Peng, et al.
Published: (2024)
Three‐Dimensional Garment Architectures for Tactile Embodied Intelligence
by: Junhua Huang, et al.
Published: (2026)
by: Junhua Huang, et al.
Published: (2026)
Cyclic Fusion of Measuring Information in Curved Elastomer Contact via Vision-Based Tactile Sensing
by: Li, Zilan, et al.
Published: (2023)
by: Li, Zilan, et al.
Published: (2023)
ReTac-ACT: A State-Gated Vision-Tactile Fusion Transformer for Precision Assembly
by: Ruan, Minchi, et al.
Published: (2026)
by: Ruan, Minchi, et al.
Published: (2026)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
UTact: Underwater Vision‐Based Tactile Sensor with Geometry Reconstruction and Contact Force Estimation
by: Qiyi Zhang, et al.
Published: (2025)
by: Qiyi Zhang, et al.
Published: (2025)
A Vision-Based Tactile Sensing System for Multimodal Contact Information Perception via Neural Network
by: Xu, Weiliang, et al.
Published: (2023)
by: Xu, Weiliang, et al.
Published: (2023)
Soft Contact Simulation and Manipulation Learning of Deformable Objects with Vision-based Tactile Sensor
by: Shan, Jianhua, et al.
Published: (2024)
by: Shan, Jianhua, et al.
Published: (2024)
A Survey on 3D Skeleton-Based Action Recognition Using Learning Method
by: Ren, Bin, et al.
Published: (2020)
by: Ren, Bin, et al.
Published: (2020)
Safety of Embodied Navigation: A Survey
by: Wang, Zixia, et al.
Published: (2025)
by: Wang, Zixia, et al.
Published: (2025)
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
by: Guan, Weifan, et al.
Published: (2025)
by: Guan, Weifan, et al.
Published: (2025)
Identifying Independent Components and Internal Process Order Parameters in Nonequilibrium Multicomponent Nonstoichiometric Compounds
by: Ji, Yanzhou, et al.
Published: (2024)
by: Ji, Yanzhou, et al.
Published: (2024)
radarODE: An ODE-Embedded Deep Learning Model for Contactless ECG Reconstruction from Millimeter-Wave Radar
by: Zhang, Yuanyuan, et al.
Published: (2024)
by: Zhang, Yuanyuan, et al.
Published: (2024)
Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
by: Xiong, Chaoran, et al.
Published: (2025)
by: Xiong, Chaoran, et al.
Published: (2025)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
by: Han, Xiaofeng, et al.
Published: (2025)
by: Han, Xiaofeng, et al.
Published: (2025)
Tactile Modality Fusion for Vision-Language-Action Models
by: Morissette, Charlotte, et al.
Published: (2026)
by: Morissette, Charlotte, et al.
Published: (2026)
Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
A Folded, Structure‐Integrated Bimodal Sensor Enabling Non‐Contact and Tactile Perception for Intelligent Robots
by: Weixiong Yang, et al.
Published: (2026)
by: Weixiong Yang, et al.
Published: (2026)
FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2024)
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2024)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
by: Zhang, Jiyao, et al.
Published: (2026)
by: Zhang, Jiyao, et al.
Published: (2026)
Multimodal Referring Segmentation: A Survey
by: Ding, Henghui, et al.
Published: (2025)
by: Ding, Henghui, et al.
Published: (2025)
Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
Flexible Electronics‐Driven Intelligent Oral Healthcare Paradigms and Next‐Generation Preventive Diagnostics
by: Hongwei Sheng, et al.
Published: (2025)
by: Hongwei Sheng, et al.
Published: (2025)
PEPL: Precision-Enhanced Pseudo-Labeling for Fine-Grained Image Classification in Semi-Supervised Learning
by: Tian, Bowen, et al.
Published: (2024)
by: Tian, Bowen, et al.
Published: (2024)
Embodied Intelligence for Flexible Manufacturing: A Survey
by: Xu, Kai, et al.
Published: (2025)
by: Xu, Kai, et al.
Published: (2025)
Rethinking Cognition: Morphological Info-Computation and the Embodied Paradigm in Life and Artificial Intelligence
by: Dodig-Crnkovic, Gordana
Published: (2024)
by: Dodig-Crnkovic, Gordana
Published: (2024)
Embodied Intelligent Spectrum Management: A New Paradigm for Dynamic Spectrum Access
by: Diao, Yihe, et al.
Published: (2026)
by: Diao, Yihe, et al.
Published: (2026)
Large Multimodal Models for Embodied Intelligent Driving: The Next Frontier in Self-Driving?
by: Zhang, Long, et al.
Published: (2026)
by: Zhang, Long, et al.
Published: (2026)
A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
by: Liu, Daizong, et al.
Published: (2024)
by: Liu, Daizong, et al.
Published: (2024)
Embodied Tactile Perception of Soft Objects Properties
by: Dutta, Anirvan, et al.
Published: (2025)
by: Dutta, Anirvan, et al.
Published: (2025)
Similar Items
-
Cognitive Disentanglement for Referring Multi-Object Tracking
by: Liang, Shaofeng, et al.
Published: (2025) -
Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
by: Guan, Runwei, et al.
Published: (2025) -
AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios
by: Chen, Chenglizhao, et al.
Published: (2025) -
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
by: Hou, Minghui, et al.
Published: (2025) -
AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild
by: Sun, Xiaolou, et al.
Published: (2026)