KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hosseinzadeh, Mehdi, Wong, King Hang, Dayoub, Feras |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
von: Wang, Wenze, et al.
Veröffentlicht: (2026)
von: Wang, Wenze, et al.
Veröffentlicht: (2026)
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
von: Podgorski, Stefan, et al.
Veröffentlicht: (2025)
von: Podgorski, Stefan, et al.
Veröffentlicht: (2025)
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
von: Garg, Sourav, et al.
Veröffentlicht: (2024)
von: Garg, Sourav, et al.
Veröffentlicht: (2024)
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
von: Abraham, Savitha Sam, et al.
Veröffentlicht: (2024)
von: Abraham, Savitha Sam, et al.
Veröffentlicht: (2024)
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
von: Holden, Lachlan, et al.
Veröffentlicht: (2026)
von: Holden, Lachlan, et al.
Veröffentlicht: (2026)
QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
von: Chapman, Nicolas Harvey, et al.
Veröffentlicht: (2025)
von: Chapman, Nicolas Harvey, et al.
Veröffentlicht: (2025)
Embodied Domain Adaptation for Object Detection
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards
von: Patel, Shivansh, et al.
Veröffentlicht: (2025)
von: Patel, Shivansh, et al.
Veröffentlicht: (2025)
ObjectReact: Learning Object-Relative Control for Visual Navigation
von: Garg, Sourav, et al.
Veröffentlicht: (2025)
von: Garg, Sourav, et al.
Veröffentlicht: (2025)
VERDI: VLM-Embedded Reasoning for Autonomous Driving
von: Feng, Bowen, et al.
Veröffentlicht: (2025)
von: Feng, Bowen, et al.
Veröffentlicht: (2025)
STAR: A Foundation Model-driven Framework for Robust Task Planning and Failure Recovery in Robotic Systems
von: Sakib, Md Sadman, et al.
Veröffentlicht: (2025)
von: Sakib, Md Sadman, et al.
Veröffentlicht: (2025)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
von: Wang, Beichen, et al.
Veröffentlicht: (2024)
von: Wang, Beichen, et al.
Veröffentlicht: (2024)
AnyTraverse: An off-road traversability framework with VLM and human operator in the loop
von: Sahu, Sattwik, et al.
Veröffentlicht: (2025)
von: Sahu, Sattwik, et al.
Veröffentlicht: (2025)
Work Zones challenge VLM Trajectory Planning: Toward Mitigation and Robust Autonomous Driving
von: Liao, Yifan, et al.
Veröffentlicht: (2025)
von: Liao, Yifan, et al.
Veröffentlicht: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
Robix: A Unified Model for Robot Interaction, Reasoning and Planning
von: Fang, Huang, et al.
Veröffentlicht: (2025)
von: Fang, Huang, et al.
Veröffentlicht: (2025)
VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving
von: Huang, Zilin, et al.
Veröffentlicht: (2024)
von: Huang, Zilin, et al.
Veröffentlicht: (2024)
DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving
von: Huang, Zilin, et al.
Veröffentlicht: (2026)
von: Huang, Zilin, et al.
Veröffentlicht: (2026)
InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects
von: Mohamed, Ibrahim Sheikh, et al.
Veröffentlicht: (2025)
von: Mohamed, Ibrahim Sheikh, et al.
Veröffentlicht: (2025)
BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2024)
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2024)
AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots
von: Zhang, Likui, et al.
Veröffentlicht: (2026)
von: Zhang, Likui, et al.
Veröffentlicht: (2026)
CurricuVLM: Towards Safe Autonomous Driving via Personalized Safety-Critical Curriculum Learning with Vision-Language Models
von: Sheng, Zihao, et al.
Veröffentlicht: (2025)
von: Sheng, Zihao, et al.
Veröffentlicht: (2025)
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
von: Jiang, Guangqi, et al.
Veröffentlicht: (2024)
von: Jiang, Guangqi, et al.
Veröffentlicht: (2024)
Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving
von: Huang, Zilin, et al.
Veröffentlicht: (2026)
von: Huang, Zilin, et al.
Veröffentlicht: (2026)
Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding
von: Tavella, Federico, et al.
Veröffentlicht: (2025)
von: Tavella, Federico, et al.
Veröffentlicht: (2025)
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
von: Atuhurra, Jesse, et al.
Veröffentlicht: (2025)
von: Atuhurra, Jesse, et al.
Veröffentlicht: (2025)
Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning
von: Zheng, Gehan, et al.
Veröffentlicht: (2026)
von: Zheng, Gehan, et al.
Veröffentlicht: (2026)
Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent
von: Yu, Che Rin, et al.
Veröffentlicht: (2025)
von: Yu, Che Rin, et al.
Veröffentlicht: (2025)
FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous Tokens
von: Zhong, Yiming, et al.
Veröffentlicht: (2025)
von: Zhong, Yiming, et al.
Veröffentlicht: (2025)
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
von: Zhou, Enshen, et al.
Veröffentlicht: (2024)
von: Zhou, Enshen, et al.
Veröffentlicht: (2024)
Visual Foresight for Robotic Stow: A Diffusion-Based World Model from Sparse Snapshots
von: Zhang, Lijun, et al.
Veröffentlicht: (2026)
von: Zhang, Lijun, et al.
Veröffentlicht: (2026)
A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
von: Jahangard, Simindokht, et al.
Veröffentlicht: (2025)
von: Jahangard, Simindokht, et al.
Veröffentlicht: (2025)
Distracted Robot: How Visual Clutter Undermine Robotic Manipulation
von: Rasouli, Amir, et al.
Veröffentlicht: (2025)
von: Rasouli, Amir, et al.
Veröffentlicht: (2025)
HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
von: Zhang, Jianke, et al.
Veröffentlicht: (2024)
von: Zhang, Jianke, et al.
Veröffentlicht: (2024)
From Failures to Fixes: LLM-Driven Scenario Repair for Self-Evolving Autonomous Driving
von: Xia, Xinyu, et al.
Veröffentlicht: (2025)
von: Xia, Xinyu, et al.
Veröffentlicht: (2025)
Your Robot Will Feel You Now: Empathy in Robots and Embodied Agents
von: Lim, Angelica, et al.
Veröffentlicht: (2026)
von: Lim, Angelica, et al.
Veröffentlicht: (2026)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities
von: Jiang, Yunfan, et al.
Veröffentlicht: (2025)
von: Jiang, Yunfan, et al.
Veröffentlicht: (2025)
Robotic Visual Instruction
von: Li, Yanbang, et al.
Veröffentlicht: (2025)
von: Li, Yanbang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
von: Wang, Wenze, et al.
Veröffentlicht: (2026) -
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
von: Podgorski, Stefan, et al.
Veröffentlicht: (2025) -
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
von: Garg, Sourav, et al.
Veröffentlicht: (2024) -
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
von: Abraham, Savitha Sam, et al.
Veröffentlicht: (2024) -
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
von: Holden, Lachlan, et al.
Veröffentlicht: (2026)