Self-Refining Vision Language Model for Robotic Failure Detection and Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Carl, Wang, Xiaojie, Yong, Silong, Sheng, Stephen, Mao, Huitan, Srinivasan, Sriram, Nambi, Manikantan, Zhang, Amy, Dattatreya, Yesh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generalizable Dense Reward for Long-Horizon Robotic Tasks
by: Yong, Silong, et al.
Published: (2026)
by: Yong, Silong, et al.
Published: (2026)
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
by: Grislain, Clemence, et al.
Published: (2025)
by: Grislain, Clemence, et al.
Published: (2025)
HydroShear: Hydroelastic Shear Simulation for Tactile Sim-to-Real Reinforcement Learning
by: Dang, An, et al.
Published: (2026)
by: Dang, An, et al.
Published: (2026)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
by: Pacaud, Paul, et al.
Published: (2025)
by: Pacaud, Paul, et al.
Published: (2025)
Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts
by: Chen, Hongyi, et al.
Published: (2024)
by: Chen, Hongyi, et al.
Published: (2024)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
Addressing Failures in Robotics using Vision-Based Language Models (VLMs) and Behavior Trees (BT)
by: Ahmad, Faseeh, et al.
Published: (2024)
by: Ahmad, Faseeh, et al.
Published: (2024)
SAFE: Multitask Failure Detection for Vision-Language-Action Models
by: Gu, Qiao, et al.
Published: (2025)
by: Gu, Qiao, et al.
Published: (2025)
Large Language Models Enable Automated Formative Feedback in Human-Robot Interaction Tasks
by: Jensen, Emily, et al.
Published: (2024)
by: Jensen, Emily, et al.
Published: (2024)
Vision-Language Model-based Physical Reasoning for Robot Liquid Perception
by: Lai, Wenqiang, et al.
Published: (2024)
by: Lai, Wenqiang, et al.
Published: (2024)
VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation
by: Zhao, Wentao, et al.
Published: (2024)
by: Zhao, Wentao, et al.
Published: (2024)
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models
by: Wu, Yanru, et al.
Published: (2026)
by: Wu, Yanru, et al.
Published: (2026)
Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models
by: Chen, Annie S., et al.
Published: (2024)
by: Chen, Annie S., et al.
Published: (2024)
Generalizable Vision-Language Few-Shot Adaptation with Predictive Prompts and Negative Learning
by: Mandalika, Sriram
Published: (2025)
by: Mandalika, Sriram
Published: (2025)
ReconVLA: An Uncertainty-Guided and Failure-Aware Vision-Language-Action Framework for Robotic Control
by: Chen, Lingling, et al.
Published: (2026)
by: Chen, Lingling, et al.
Published: (2026)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
by: Huang, Yuzhi, et al.
Published: (2026)
by: Huang, Yuzhi, et al.
Published: (2026)
Graph-Fused Vision-Language-Action for Policy Reasoning in Multi-Arm Robotic Manipulation
by: Li, Shunlei, et al.
Published: (2025)
by: Li, Shunlei, et al.
Published: (2025)
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
by: Peng, Zhenghao "Mark", et al.
Published: (2025)
by: Peng, Zhenghao "Mark", et al.
Published: (2025)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
by: Han, Yi, et al.
Published: (2025)
by: Han, Yi, et al.
Published: (2025)
DyNaVLM: Zero-Shot Vision-Language Navigation System with Dynamic Viewpoints and Self-Refining Graph Memory
by: Ji, Zihe, et al.
Published: (2025)
by: Ji, Zihe, et al.
Published: (2025)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
Latent Space Reinforcement Learning for Multi-Robot Exploration
by: Rajasekar, Sriram, et al.
Published: (2026)
by: Rajasekar, Sriram, et al.
Published: (2026)
Multi-Robot Object SLAM Using Distributed Variational Inference
by: Cao, Hanwen, et al.
Published: (2024)
by: Cao, Hanwen, et al.
Published: (2024)
Experiences from Benchmarking Vision-Language-Action Models for Robotic Manipulation
by: Zhang, Yihao, et al.
Published: (2025)
by: Zhang, Yihao, et al.
Published: (2025)
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
by: Zhou, Enshen, et al.
Published: (2024)
by: Zhou, Enshen, et al.
Published: (2024)
MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction
by: Deigmoeller, Joerg, et al.
Published: (2026)
by: Deigmoeller, Joerg, et al.
Published: (2026)
Decentralized LLM-Driven Coordination of Acoustic Robots for Contactless Object Manipulation
by: Wang, Yingying, et al.
Published: (2026)
by: Wang, Yingying, et al.
Published: (2026)
SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model
by: Zhu, Shaoting, et al.
Published: (2024)
by: Zhu, Shaoting, et al.
Published: (2024)
Vision-Based Reasoning with Topology-Encoded Graphs for Anatomical Path Disambiguation in Robot-Assisted Endovascular Navigation
by: Zhao, Jiyuan, et al.
Published: (2026)
by: Zhao, Jiyuan, et al.
Published: (2026)
ViLAM: Distilling Vision-Language Reasoning into Attention Maps for Social Robot Navigation
by: Elnoor, Mohamed, et al.
Published: (2025)
by: Elnoor, Mohamed, et al.
Published: (2025)
Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment
by: Liu, Kangcheng, et al.
Published: (2023)
by: Liu, Kangcheng, et al.
Published: (2023)
Empathetic Motion Generation for Humanoid Educational Robots via Reasoning-Guided Vision--Language--Motion Diffusion Architecture
by: Sun, Fuze, et al.
Published: (2026)
by: Sun, Fuze, et al.
Published: (2026)
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
by: Guo, Wenxuan, et al.
Published: (2026)
by: Guo, Wenxuan, et al.
Published: (2026)
MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
by: Huang, Wenhui, et al.
Published: (2025)
by: Huang, Wenhui, et al.
Published: (2025)
Failing Forward: Adaptive Failure-Informed Learning for Vision-Language-Action Models
by: Zheng, Meng, et al.
Published: (2026)
by: Zheng, Meng, et al.
Published: (2026)
ERR@HRI 2024 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Interactions
by: Spitale, Micol, et al.
Published: (2024)
by: Spitale, Micol, et al.
Published: (2024)
Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention
by: Mao, Yanbo, et al.
Published: (2025)
by: Mao, Yanbo, et al.
Published: (2025)
LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
by: Hong, Youngjin, et al.
Published: (2025)
by: Hong, Youngjin, et al.
Published: (2025)
Similar Items
-
Generalizable Dense Reward for Long-Horizon Robotic Tasks
by: Yong, Silong, et al.
Published: (2026) -
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
by: Duan, Jiafei, et al.
Published: (2024) -
I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
by: Grislain, Clemence, et al.
Published: (2025) -
HydroShear: Hydroelastic Shear Simulation for Tactile Sim-to-Real Reinforcement Learning
by: Dang, An, et al.
Published: (2026) -
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
by: Pacaud, Paul, et al.
Published: (2025)