Gespeichert in:
| Hauptverfasser: | Zhou, Yiyun, Xu, Mingjing, Shi, Jingwei, Li, Quanjiang, Chen, Jingyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2511.11512 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning
von: Zhou, Yiyun, et al.
Veröffentlicht: (2026)
von: Zhou, Yiyun, et al.
Veröffentlicht: (2026)
Tactile Modality Fusion for Vision-Language-Action Models
von: Morissette, Charlotte, et al.
Veröffentlicht: (2026)
von: Morissette, Charlotte, et al.
Veröffentlicht: (2026)
Binding Touch to Everything: Learning Unified Multimodal Tactile Representations
von: Yang, Fengyu, et al.
Veröffentlicht: (2024)
von: Yang, Fengyu, et al.
Veröffentlicht: (2024)
AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception
von: Feng, Ruoxuan, et al.
Veröffentlicht: (2026)
von: Feng, Ruoxuan, et al.
Veröffentlicht: (2026)
VitaTouch: Property-Aware Vision-Tactile-Language Model for Robotic Quality Inspection in Manufacturing
von: Zong, Junyi, et al.
Veröffentlicht: (2026)
von: Zong, Junyi, et al.
Veröffentlicht: (2026)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
von: Lin, Tao, et al.
Veröffentlicht: (2025)
von: Lin, Tao, et al.
Veröffentlicht: (2025)
Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment
von: Liu, Kangcheng, et al.
Veröffentlicht: (2023)
von: Liu, Kangcheng, et al.
Veröffentlicht: (2023)
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
von: Su, Xia, et al.
Veröffentlicht: (2026)
von: Su, Xia, et al.
Veröffentlicht: (2026)
Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
von: Li, Yuyang, et al.
Veröffentlicht: (2025)
von: Li, Yuyang, et al.
Veröffentlicht: (2025)
FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors
von: Shahidzadeh, Amir-Hossein, et al.
Veröffentlicht: (2024)
von: Shahidzadeh, Amir-Hossein, et al.
Veröffentlicht: (2024)
Sensor-Invariant Tactile Representation
von: Gupta, Harsh, et al.
Veröffentlicht: (2025)
von: Gupta, Harsh, et al.
Veröffentlicht: (2025)
Bridging Text and Vision: A Multi-View Text-Vision Registration Approach for Cross-Modal Place Recognition
von: Shang, Tianyi, et al.
Veröffentlicht: (2025)
von: Shang, Tianyi, et al.
Veröffentlicht: (2025)
A Touch, Vision, and Language Dataset for Multimodal Alignment
von: Fu, Letian, et al.
Veröffentlicht: (2024)
von: Fu, Letian, et al.
Veröffentlicht: (2024)
Does Peer Observation Help? Vision-Sharing Collaboration for Vision-Language Navigation
von: Jin, Qunchao, et al.
Veröffentlicht: (2026)
von: Jin, Qunchao, et al.
Veröffentlicht: (2026)
ConViTac: Aligning Visual-Tactile Fusion with Contrastive Representations
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment
von: Singh, Binod, et al.
Veröffentlicht: (2025)
von: Singh, Binod, et al.
Veröffentlicht: (2025)
TLA: Tactile-Language-Action Model for Contact-Rich Manipulation
von: Hao, Peng, et al.
Veröffentlicht: (2025)
von: Hao, Peng, et al.
Veröffentlicht: (2025)
Taccel: Scaling Up Vision-based Tactile Robotics via High-performance GPU Simulation
von: Li, Yuyang, et al.
Veröffentlicht: (2025)
von: Li, Yuyang, et al.
Veröffentlicht: (2025)
SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
von: Renz, Katrin, et al.
Veröffentlicht: (2025)
von: Renz, Katrin, et al.
Veröffentlicht: (2025)
Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks
von: Zhao, Jialiang, et al.
Veröffentlicht: (2024)
von: Zhao, Jialiang, et al.
Veröffentlicht: (2024)
Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
Tacchi 2.0: A Low Computational Cost and Comprehensive Dynamic Contact Simulator for Vision-based Tactile Sensors
von: Sun, Yuhao, et al.
Veröffentlicht: (2025)
von: Sun, Yuhao, et al.
Veröffentlicht: (2025)
GelBelt: A Vision-based Tactile Sensor for Continuous Sensing of Large Surfaces
von: Mirzaee, Mohammad Amin, et al.
Veröffentlicht: (2025)
von: Mirzaee, Mohammad Amin, et al.
Veröffentlicht: (2025)
Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms
von: Cao, Zhixiang, et al.
Veröffentlicht: (2026)
von: Cao, Zhixiang, et al.
Veröffentlicht: (2026)
Ensuring Force Safety in Vision-Guided Robotic Manipulation via Implicit Tactile Calibration
von: Wei, Lai, et al.
Veröffentlicht: (2024)
von: Wei, Lai, et al.
Veröffentlicht: (2024)
Multi-Modal UAV Detection, Classification and Tracking Algorithm -- Technical Report for CVPR 2024 UG2 Challenge
von: Deng, Tianchen, et al.
Veröffentlicht: (2024)
von: Deng, Tianchen, et al.
Veröffentlicht: (2024)
ReViP: Mitigating False Completion in Vision-Language-Action Models with Vision-Proprioception Rebalance
von: Li, Zhuohao, et al.
Veröffentlicht: (2026)
von: Li, Zhuohao, et al.
Veröffentlicht: (2026)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025)
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
von: Li, Haoyuan, et al.
Veröffentlicht: (2026)
von: Li, Haoyuan, et al.
Veröffentlicht: (2026)
AffordanceLLM: Grounding Affordance from Vision Language Models
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
von: Guo, Jianing, et al.
Veröffentlicht: (2025)
von: Guo, Jianing, et al.
Veröffentlicht: (2025)
GARAD-SLAM: 3D GAussian splatting for Real-time Anti Dynamic SLAM
von: Li, Mingrui, et al.
Veröffentlicht: (2025)
von: Li, Mingrui, et al.
Veröffentlicht: (2025)
Vision-Only Gaussian Splatting for Collaborative Semantic Occupancy Prediction
von: Chen, Cheng, et al.
Veröffentlicht: (2025)
von: Chen, Cheng, et al.
Veröffentlicht: (2025)
Symmetry-Aware Fusion of Vision and Tactile Sensing via Bilateral Force Priors for Robotic Manipulation
von: Lee, Wonju, et al.
Veröffentlicht: (2026)
von: Lee, Wonju, et al.
Veröffentlicht: (2026)
DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
von: Shi, Haoxiang, et al.
Veröffentlicht: (2025)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2025)
ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning
von: Tian, Yufeng, et al.
Veröffentlicht: (2026)
von: Tian, Yufeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning
von: Zhou, Yiyun, et al.
Veröffentlicht: (2026) -
Tactile Modality Fusion for Vision-Language-Action Models
von: Morissette, Charlotte, et al.
Veröffentlicht: (2026) -
Binding Touch to Everything: Learning Unified Multimodal Tactile Representations
von: Yang, Fengyu, et al.
Veröffentlicht: (2024) -
AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception
von: Feng, Ruoxuan, et al.
Veröffentlicht: (2026) -
VitaTouch: Property-Aware Vision-Tactile-Language Model for Robotic Quality Inspection in Manufacturing
von: Zong, Junyi, et al.
Veröffentlicht: (2026)