KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
Fuente:
arXiv
Salvato in:
| Autori principali: | Hosseinzadeh, Mehdi, Wong, King Hang, Dayoub, Feras |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
di: Wang, Wenze, et al.
Pubblicazione: (2026)
di: Wang, Wenze, et al.
Pubblicazione: (2026)
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
di: Podgorski, Stefan, et al.
Pubblicazione: (2025)
di: Podgorski, Stefan, et al.
Pubblicazione: (2025)
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
di: Garg, Sourav, et al.
Pubblicazione: (2024)
di: Garg, Sourav, et al.
Pubblicazione: (2024)
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
di: Abraham, Savitha Sam, et al.
Pubblicazione: (2024)
di: Abraham, Savitha Sam, et al.
Pubblicazione: (2024)
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
di: Holden, Lachlan, et al.
Pubblicazione: (2026)
di: Holden, Lachlan, et al.
Pubblicazione: (2026)
QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
di: Chapman, Nicolas Harvey, et al.
Pubblicazione: (2025)
di: Chapman, Nicolas Harvey, et al.
Pubblicazione: (2025)
Embodied Domain Adaptation for Object Detection
di: Shi, Xiangyu, et al.
Pubblicazione: (2025)
di: Shi, Xiangyu, et al.
Pubblicazione: (2025)
A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards
di: Patel, Shivansh, et al.
Pubblicazione: (2025)
di: Patel, Shivansh, et al.
Pubblicazione: (2025)
ObjectReact: Learning Object-Relative Control for Visual Navigation
di: Garg, Sourav, et al.
Pubblicazione: (2025)
di: Garg, Sourav, et al.
Pubblicazione: (2025)
VERDI: VLM-Embedded Reasoning for Autonomous Driving
di: Feng, Bowen, et al.
Pubblicazione: (2025)
di: Feng, Bowen, et al.
Pubblicazione: (2025)
STAR: A Foundation Model-driven Framework for Robust Task Planning and Failure Recovery in Robotic Systems
di: Sakib, Md Sadman, et al.
Pubblicazione: (2025)
di: Sakib, Md Sadman, et al.
Pubblicazione: (2025)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
di: Wang, Beichen, et al.
Pubblicazione: (2024)
di: Wang, Beichen, et al.
Pubblicazione: (2024)
AnyTraverse: An off-road traversability framework with VLM and human operator in the loop
di: Sahu, Sattwik, et al.
Pubblicazione: (2025)
di: Sahu, Sattwik, et al.
Pubblicazione: (2025)
Work Zones challenge VLM Trajectory Planning: Toward Mitigation and Robust Autonomous Driving
di: Liao, Yifan, et al.
Pubblicazione: (2025)
di: Liao, Yifan, et al.
Pubblicazione: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
di: Kim, Dongwon, et al.
Pubblicazione: (2026)
di: Kim, Dongwon, et al.
Pubblicazione: (2026)
Robix: A Unified Model for Robot Interaction, Reasoning and Planning
di: Fang, Huang, et al.
Pubblicazione: (2025)
di: Fang, Huang, et al.
Pubblicazione: (2025)
VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving
di: Huang, Zilin, et al.
Pubblicazione: (2024)
di: Huang, Zilin, et al.
Pubblicazione: (2024)
DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving
di: Huang, Zilin, et al.
Pubblicazione: (2026)
di: Huang, Zilin, et al.
Pubblicazione: (2026)
InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects
di: Mohamed, Ibrahim Sheikh, et al.
Pubblicazione: (2025)
di: Mohamed, Ibrahim Sheikh, et al.
Pubblicazione: (2025)
BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment
di: Hosseinzadeh, Mehdi, et al.
Pubblicazione: (2024)
di: Hosseinzadeh, Mehdi, et al.
Pubblicazione: (2024)
AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots
di: Zhang, Likui, et al.
Pubblicazione: (2026)
di: Zhang, Likui, et al.
Pubblicazione: (2026)
CurricuVLM: Towards Safe Autonomous Driving via Personalized Safety-Critical Curriculum Learning with Vision-Language Models
di: Sheng, Zihao, et al.
Pubblicazione: (2025)
di: Sheng, Zihao, et al.
Pubblicazione: (2025)
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
di: Shi, Xiangyu, et al.
Pubblicazione: (2025)
di: Shi, Xiangyu, et al.
Pubblicazione: (2025)
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
di: Jiang, Guangqi, et al.
Pubblicazione: (2024)
di: Jiang, Guangqi, et al.
Pubblicazione: (2024)
Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving
di: Huang, Zilin, et al.
Pubblicazione: (2026)
di: Huang, Zilin, et al.
Pubblicazione: (2026)
Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding
di: Tavella, Federico, et al.
Pubblicazione: (2025)
di: Tavella, Federico, et al.
Pubblicazione: (2025)
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
di: Atuhurra, Jesse, et al.
Pubblicazione: (2025)
di: Atuhurra, Jesse, et al.
Pubblicazione: (2025)
Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning
di: Zheng, Gehan, et al.
Pubblicazione: (2026)
di: Zheng, Gehan, et al.
Pubblicazione: (2026)
Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent
di: Yu, Che Rin, et al.
Pubblicazione: (2025)
di: Yu, Che Rin, et al.
Pubblicazione: (2025)
FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous Tokens
di: Zhong, Yiming, et al.
Pubblicazione: (2025)
di: Zhong, Yiming, et al.
Pubblicazione: (2025)
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
di: Zhou, Enshen, et al.
Pubblicazione: (2024)
di: Zhou, Enshen, et al.
Pubblicazione: (2024)
Visual Foresight for Robotic Stow: A Diffusion-Based World Model from Sparse Snapshots
di: Zhang, Lijun, et al.
Pubblicazione: (2026)
di: Zhang, Lijun, et al.
Pubblicazione: (2026)
A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
di: Jahangard, Simindokht, et al.
Pubblicazione: (2025)
di: Jahangard, Simindokht, et al.
Pubblicazione: (2025)
Distracted Robot: How Visual Clutter Undermine Robotic Manipulation
di: Rasouli, Amir, et al.
Pubblicazione: (2025)
di: Rasouli, Amir, et al.
Pubblicazione: (2025)
HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
di: Zhang, Jianke, et al.
Pubblicazione: (2024)
di: Zhang, Jianke, et al.
Pubblicazione: (2024)
From Failures to Fixes: LLM-Driven Scenario Repair for Self-Evolving Autonomous Driving
di: Xia, Xinyu, et al.
Pubblicazione: (2025)
di: Xia, Xinyu, et al.
Pubblicazione: (2025)
Your Robot Will Feel You Now: Empathy in Robots and Embodied Agents
di: Lim, Angelica, et al.
Pubblicazione: (2026)
di: Lim, Angelica, et al.
Pubblicazione: (2026)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities
di: Jiang, Yunfan, et al.
Pubblicazione: (2025)
di: Jiang, Yunfan, et al.
Pubblicazione: (2025)
Robotic Visual Instruction
di: Li, Yanbang, et al.
Pubblicazione: (2025)
di: Li, Yanbang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
di: Wang, Wenze, et al.
Pubblicazione: (2026) -
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
di: Podgorski, Stefan, et al.
Pubblicazione: (2025) -
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
di: Garg, Sourav, et al.
Pubblicazione: (2024) -
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
di: Abraham, Savitha Sam, et al.
Pubblicazione: (2024) -
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
di: Holden, Lachlan, et al.
Pubblicazione: (2026)