AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Jiafei, Pumacay, Wilbert, Kumar, Nishanth, Wang, Yi Ru, Tian, Shulin, Yuan, Wentao, Krishna, Ranjay, Fox, Dieter, Mandlekar, Ajay, Guo, Yijie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
by: Pumacay, Wilbert, et al.
Published: (2024)
by: Pumacay, Wilbert, et al.
Published: (2024)
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
by: Fang, Haoquan, et al.
Published: (2025)
by: Fang, Haoquan, et al.
Published: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
by: Wang, Yi Ru, et al.
Published: (2025)
by: Wang, Yi Ru, et al.
Published: (2025)
EVE: Enabling Anyone to Train Robots using Augmented Reality
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
by: Tur, Yalcin, et al.
Published: (2026)
by: Tur, Yalcin, et al.
Published: (2026)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
SPIRE: Synergistic Planning, Imitation, and Reinforcement Learning for Long-Horizon Manipulation
by: Zhou, Zihan, et al.
Published: (2024)
by: Zhou, Zihan, et al.
Published: (2024)
MolmoAct: Action Reasoning Models that can Reason in Space
by: Lee, Jason, et al.
Published: (2025)
by: Lee, Jason, et al.
Published: (2025)
IntervenGen: Interventional Data Generation for Robust and Data-Efficient Robot Imitation Learning
by: Hoque, Ryan, et al.
Published: (2024)
by: Hoque, Ryan, et al.
Published: (2024)
DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation
by: Mandi, Zhao, et al.
Published: (2025)
by: Mandi, Zhao, et al.
Published: (2025)
SkillMimicGen: Automated Demonstration Generation for Efficient Skill Learning and Deployment
by: Garrett, Caelan, et al.
Published: (2024)
by: Garrett, Caelan, et al.
Published: (2024)
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
by: Chen, Shirui, et al.
Published: (2026)
by: Chen, Shirui, et al.
Published: (2026)
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
by: Kim, Yejin, et al.
Published: (2026)
by: Kim, Yejin, et al.
Published: (2026)
MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
by: Deshpande, Abhay, et al.
Published: (2026)
by: Deshpande, Abhay, et al.
Published: (2026)
SoftMimicGen: A Data Generation System for Scalable Robot Learning in Deformable Object Manipulation
by: Moghani, Masoud, et al.
Published: (2026)
by: Moghani, Masoud, et al.
Published: (2026)
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
by: Cheng, Long, et al.
Published: (2025)
by: Cheng, Long, et al.
Published: (2025)
TeleopLab: Accessible and Intuitive Teleoperation of a Robotic Manipulator for Remote Labs
by: Chen, Ziling, et al.
Published: (2025)
by: Chen, Ziling, et al.
Published: (2025)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
by: Eftekhar, Ainaz, et al.
Published: (2023)
by: Eftekhar, Ainaz, et al.
Published: (2023)
ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Point Bridge: 3D Representations for Cross Domain Policy Learning
by: Haldar, Siddhant, et al.
Published: (2026)
by: Haldar, Siddhant, et al.
Published: (2026)
vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models
by: Choi, Suhwan, et al.
Published: (2026)
by: Choi, Suhwan, et al.
Published: (2026)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
by: Pacaud, Paul, et al.
Published: (2025)
by: Pacaud, Paul, et al.
Published: (2025)
RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields
by: Sagar, Som, et al.
Published: (2024)
by: Sagar, Som, et al.
Published: (2024)
What Matters in Learning from Large-Scale Datasets for Robot Manipulation
by: Saxena, Vaibhav, et al.
Published: (2025)
by: Saxena, Vaibhav, et al.
Published: (2025)
Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation
by: Maddukuri, Abhiram, et al.
Published: (2025)
by: Maddukuri, Abhiram, et al.
Published: (2025)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks
by: Grotz, Markus, et al.
Published: (2024)
by: Grotz, Markus, et al.
Published: (2024)
RVT-2: Learning Precise Manipulation from Few Demonstrations
by: Goyal, Ankit, et al.
Published: (2024)
by: Goyal, Ankit, et al.
Published: (2024)
Open-World Task and Motion Planning via Vision-Language Model Generated Constraints
by: Kumar, Nishanth, et al.
Published: (2024)
by: Kumar, Nishanth, et al.
Published: (2024)
Enhancing Power Quality in Smart Grid‐Connected Renewable Energy Systems Using A Hybrid Deep Learning‐Based DSTATCOM: Self‐Improved Jellyfish Optimizer (Si‐Jo)
by: K. Nishanth, et al.
Published: (2024)
by: K. Nishanth, et al.
Published: (2024)
ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning
by: Zhou, Zihan, et al.
Published: (2025)
by: Zhou, Zihan, et al.
Published: (2025)
NOD-TAMP: Generalizable Long-Horizon Planning with Neural Object Descriptors
by: Cheng, Shuo, et al.
Published: (2023)
by: Cheng, Shuo, et al.
Published: (2023)
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024)
by: Zheng, Chenhao, et al.
Published: (2024)
AHA! RSK
by: Stern, Eugene
Published: (2026)
by: Stern, Eugene
Published: (2026)
SRSA: Skill Retrieval and Adaptation for Robotic Assembly Tasks
by: Guo, Yijie, et al.
Published: (2025)
by: Guo, Yijie, et al.
Published: (2025)
ASID: Active Exploration for System Identification in Robotic Manipulation
by: Memmel, Marius, et al.
Published: (2024)
by: Memmel, Marius, et al.
Published: (2024)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
by: Kamath, Amita, et al.
Published: (2026)
by: Kamath, Amita, et al.
Published: (2026)
Similar Items
-
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
by: Duan, Jiafei, et al.
Published: (2024) -
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
by: Pumacay, Wilbert, et al.
Published: (2024) -
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
by: Fang, Haoquan, et al.
Published: (2025) -
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024) -
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
by: Lin, Zijun, et al.
Published: (2025)