Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Jimin, Jang, Huiwon, Koo, Myungkyu, Park, Jungwoo, Shin, Jinwoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
von: Won, John, et al.
Veröffentlicht: (2025)
von: Won, John, et al.
Veröffentlicht: (2025)
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
von: Koo, Myungkyu, et al.
Veröffentlicht: (2025)
von: Koo, Myungkyu, et al.
Veröffentlicht: (2025)
ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
von: Jang, Huiwon, et al.
Veröffentlicht: (2025)
von: Jang, Huiwon, et al.
Veröffentlicht: (2025)
Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
von: Kim, Dongyoung, et al.
Veröffentlicht: (2025)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2025)
Verifier-free Test-Time Sampling for Vision Language Action Models
von: Jang, Suhyeok, et al.
Veröffentlicht: (2025)
von: Jang, Suhyeok, et al.
Veröffentlicht: (2025)
Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
von: Park, Jeongeun, et al.
Veröffentlicht: (2025)
von: Park, Jeongeun, et al.
Veröffentlicht: (2025)
Visual Representation Learning with Stochastic Frame Prediction
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)
RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
von: Kim, Dongyoung, et al.
Veröffentlicht: (2026)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2026)
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
von: Park, Minho, et al.
Veröffentlicht: (2025)
von: Park, Minho, et al.
Veröffentlicht: (2025)
Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
von: Lee, Sangoh, et al.
Veröffentlicht: (2025)
von: Lee, Sangoh, et al.
Veröffentlicht: (2025)
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
Adversarial Robustification via Text-to-Image Diffusion Models
von: Choi, Daewon, et al.
Veröffentlicht: (2024)
von: Choi, Daewon, et al.
Veröffentlicht: (2024)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
von: Peng, Xiongfeng, et al.
Veröffentlicht: (2026)
von: Peng, Xiongfeng, et al.
Veröffentlicht: (2026)
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
von: Li, Xiaoqi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2026)
RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
von: Koo, Jiyeon, et al.
Veröffentlicht: (2025)
von: Koo, Jiyeon, et al.
Veröffentlicht: (2025)
PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments
von: Kim, Kihyun, et al.
Veröffentlicht: (2026)
von: Kim, Kihyun, et al.
Veröffentlicht: (2026)
VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
von: Bi, Jianxin, et al.
Veröffentlicht: (2025)
von: Bi, Jianxin, et al.
Veröffentlicht: (2025)
Mesh-based Photorealistic and Real-time 3D Mapping for Robust Visual Perception of Autonomous Underwater Vehicle
von: Lee, Jungwoo, et al.
Veröffentlicht: (2024)
von: Lee, Jungwoo, et al.
Veröffentlicht: (2024)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
von: Kim, Seungku, et al.
Veröffentlicht: (2026)
von: Kim, Seungku, et al.
Veröffentlicht: (2026)
Salience-guided Ground Factor for Robust Localization of Delivery Robots in Complex Urban Environments
von: Park, Jooyong, et al.
Veröffentlicht: (2024)
von: Park, Jooyong, et al.
Veröffentlicht: (2024)
RedVLA: Physical Red Teaming for Vision-Language-Action Models
von: Zhang, Yuhao, et al.
Veröffentlicht: (2026)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2026)
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling
von: Luu, Tung M., et al.
Veröffentlicht: (2025)
von: Luu, Tung M., et al.
Veröffentlicht: (2025)
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
von: Liang, Yuanchang, et al.
Veröffentlicht: (2026)
von: Liang, Yuanchang, et al.
Veröffentlicht: (2026)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
von: Zhong, Yifan, et al.
Veröffentlicht: (2025)
von: Zhong, Yifan, et al.
Veröffentlicht: (2025)
Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment
von: Wulff, Theodor, et al.
Veröffentlicht: (2026)
von: Wulff, Theodor, et al.
Veröffentlicht: (2026)
DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI
von: Yu, En, et al.
Veröffentlicht: (2026)
von: Yu, En, et al.
Veröffentlicht: (2026)
Action Hallucination in Generative Vision-Language-Action Models
von: Soh, Harold, et al.
Veröffentlicht: (2026)
von: Soh, Harold, et al.
Veröffentlicht: (2026)
Understanding Physical Properties of Unseen Deformable Objects by Leveraging Large Language Models and Robot Actions
von: Park, Changmin, et al.
Veröffentlicht: (2025)
von: Park, Changmin, et al.
Veröffentlicht: (2025)
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model
von: Guo, Yanjiang, et al.
Veröffentlicht: (2026)
von: Guo, Yanjiang, et al.
Veröffentlicht: (2026)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
von: Deng, Shengliang, et al.
Veröffentlicht: (2025)
von: Deng, Shengliang, et al.
Veröffentlicht: (2025)
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
von: Li, Zuolei, et al.
Veröffentlicht: (2025)
von: Li, Zuolei, et al.
Veröffentlicht: (2025)
A Vision-Language-Action Model for Adaptive Ultrasound-Guided Needle Insertion and Needle Tracking
von: Zhang, Yuelin, et al.
Veröffentlicht: (2026)
von: Zhang, Yuelin, et al.
Veröffentlicht: (2026)
Observing and Controlling Features in Vision-Language-Action Models
von: Buurmeijer, Hugo, et al.
Veröffentlicht: (2026)
von: Buurmeijer, Hugo, et al.
Veröffentlicht: (2026)
Embodiment Transfer Learning for Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
GraspCorrect: Robotic Grasp Correction via Vision-Language Model-Guided Feedback
von: Lee, Sungjae, et al.
Veröffentlicht: (2025)
von: Lee, Sungjae, et al.
Veröffentlicht: (2025)
Test-Time Training for Visual Foresight Vision-Language-Action Models
von: Park, Sangwu, et al.
Veröffentlicht: (2026)
von: Park, Sangwu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
von: Won, John, et al.
Veröffentlicht: (2025) -
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
von: Koo, Myungkyu, et al.
Veröffentlicht: (2025) -
ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
von: Jang, Huiwon, et al.
Veröffentlicht: (2025) -
Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
von: Kim, Dongyoung, et al.
Veröffentlicht: (2025) -
Verifier-free Test-Time Sampling for Vision Language Action Models
von: Jang, Suhyeok, et al.
Veröffentlicht: (2025)