Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yilin, Li, Anqi, Hermans, Tucker, Ramos, Fabio, Bajcsy, Andrea, Pérez-D'Arpino, Claudia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment
by: Wu, Yilin, et al.
Published: (2025)
by: Wu, Yilin, et al.
Published: (2025)
When to Act, Ask, or Learn: Uncertainty-Aware Policy Steering
by: Yuan, Jessie, et al.
Published: (2026)
by: Yuan, Jessie, et al.
Published: (2026)
Fast Explicit-Input Assistance for Teleoperation in Clutter
by: Walker, Nick, et al.
Published: (2024)
by: Walker, Nick, et al.
Published: (2024)
What You Don't Know Can Hurt You: How Well do Latent Safety Filters Understand Partially Observable Safety Constraints?
by: Kim, Matthew, et al.
Published: (2025)
by: Kim, Matthew, et al.
Published: (2025)
Inference-Time Policy Steering through Human Interactions
by: Wang, Yanwei, et al.
Published: (2024)
by: Wang, Yanwei, et al.
Published: (2024)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
by: Tian, Ran, et al.
Published: (2023)
by: Tian, Ran, et al.
Published: (2023)
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
by: Gao, Tian, et al.
Published: (2026)
by: Gao, Tian, et al.
Published: (2026)
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
by: Hsieh, Wen-Han, et al.
Published: (2025)
by: Hsieh, Wen-Han, et al.
Published: (2025)
Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment
by: Wulff, Theodor, et al.
Published: (2026)
by: Wulff, Theodor, et al.
Published: (2026)
Position: Good Embodied Reward Models Need Bad Behavior Data
by: Tian, Ran, et al.
Published: (2026)
by: Tian, Ran, et al.
Published: (2026)
Conformalized Teleoperation: Confidently Mapping Human Inputs to High-Dimensional Robot Actions
by: Zhao, Michelle, et al.
Published: (2024)
by: Zhao, Michelle, et al.
Published: (2024)
Causal Scene Narration with Runtime Safety Supervision for Vision-Language-Action Driving
by: Li, Yun, et al.
Published: (2026)
by: Li, Yun, et al.
Published: (2026)
10 Open Challenges Steering the Future of Vision-Language-Action Models
by: Poria, Soujanya, et al.
Published: (2025)
by: Poria, Soujanya, et al.
Published: (2025)
Continuous Reasoning for Vision-Language-Action
by: Wu, Yueh-Hua, et al.
Published: (2026)
by: Wu, Yueh-Hua, et al.
Published: (2026)
Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment
by: Kwok, Jacky, et al.
Published: (2026)
by: Kwok, Jacky, et al.
Published: (2026)
Towards Backdoor-Based Ownership Verification for Vision-Language-Action Models
by: Sun, Ming, et al.
Published: (2026)
by: Sun, Ming, et al.
Published: (2026)
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
by: Xu, Haiweng, et al.
Published: (2026)
by: Xu, Haiweng, et al.
Published: (2026)
Opportunities for Policy Progress: The Role of Schools in Minimizing, Mitigating, or Perpetuating Weight‐Based Stigma
by: Samantha Turner, et al.
Published: (2025)
by: Samantha Turner, et al.
Published: (2025)
23 DoF Grasping Policies from a Raw Point Cloud
by: Matak, Martin, et al.
Published: (2024)
by: Matak, Martin, et al.
Published: (2024)
InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
by: Chen, William, et al.
Published: (2026)
by: Chen, William, et al.
Published: (2026)
Sigma: The Key for Vision-Language-Action Models toward Telepathic Alignment
by: Wang, Libo
Published: (2025)
by: Wang, Libo
Published: (2025)
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
by: Yang, Siyuan, et al.
Published: (2025)
by: Yang, Siyuan, et al.
Published: (2025)
AnySafe: Adapting Latent Safety Filters at Runtime via Safety Constraint Parameterization in the Latent Space
by: Agrawal, Sankalp, et al.
Published: (2025)
by: Agrawal, Sankalp, et al.
Published: (2025)
Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment
by: Tian, Ran, et al.
Published: (2024)
by: Tian, Ran, et al.
Published: (2024)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Action Hallucination in Generative Vision-Language-Action Models
by: Soh, Harold, et al.
Published: (2026)
by: Soh, Harold, et al.
Published: (2026)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
by: Liang, Yuanchang, et al.
Published: (2026)
by: Liang, Yuanchang, et al.
Published: (2026)
Boosting Vision-Language-Action Finetuning with Feasible Action Neighborhood Prior
by: Niu, Haochen, et al.
Published: (2026)
by: Niu, Haochen, et al.
Published: (2026)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
by: Zhong, Yifan, et al.
Published: (2025)
by: Zhong, Yifan, et al.
Published: (2025)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
by: Zhong, Linqing, et al.
Published: (2026)
by: Zhong, Linqing, et al.
Published: (2026)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
by: Peng, Zhenghao "Mark", et al.
Published: (2025)
by: Peng, Zhenghao "Mark", et al.
Published: (2025)
Planning under Uncertainty to Goal Distributions
by: Conkey, Adam, et al.
Published: (2020)
by: Conkey, Adam, et al.
Published: (2020)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
by: Darabi, Nastaran, et al.
Published: (2026)
by: Darabi, Nastaran, et al.
Published: (2026)
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
by: Li, Zuolei, et al.
Published: (2025)
by: Li, Zuolei, et al.
Published: (2025)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
Similar Items
-
From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment
by: Wu, Yilin, et al.
Published: (2025) -
When to Act, Ask, or Learn: Uncertainty-Aware Policy Steering
by: Yuan, Jessie, et al.
Published: (2026) -
Fast Explicit-Input Assistance for Teleoperation in Clutter
by: Walker, Nick, et al.
Published: (2024) -
What You Don't Know Can Hurt You: How Well do Latent Safety Filters Understand Partially Observable Safety Constraints?
by: Kim, Matthew, et al.
Published: (2025) -
Inference-Time Policy Steering through Human Interactions
by: Wang, Yanwei, et al.
Published: (2024)