A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Zixin, Chen, Kanghao, Wang, Hanqing, Zhang, Hongfei, Chen, Harold Haodong, Liao, Chenfei, Guo, Litao, Chen, Ying-Cong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918250192830464
author Zhang, Zixin
Chen, Kanghao
Wang, Hanqing
Zhang, Hongfei
Chen, Harold Haodong
Liao, Chenfei
Guo, Litao
Chen, Ying-Cong
author_facet Zhang, Zixin
Chen, Kanghao
Wang, Hanqing
Zhang, Hongfei
Chen, Harold Haodong
Liao, Chenfei
Guo, Litao
Chen, Ying-Cong
contents Affordance prediction, which identifies interaction regions on objects based on language instructions, is critical for embodied AI. Prevailing end-to-end models couple high-level reasoning and low-level grounding into a single monolithic pipeline and rely on training over annotated datasets, which leads to poor generalization on novel objects and unseen environments. In this paper, we move beyond this paradigm by proposing A4-Agent, a training-free agentic framework that decouples affordance prediction into a three-stage pipeline. Our framework coordinates specialized foundation models at test time: (1) a $\textbf{Dreamer}$ that employs generative models to visualize $\textit{how}$ an interaction would look; (2) a $\textbf{Thinker}$ that utilizes large vision-language models to decide $\textit{what}$ object part to interact with; and (3) a $\textbf{Spotter}$ that orchestrates vision foundation models to precisely locate $\textit{where}$ the interaction area is. By leveraging the complementary strengths of pre-trained models without any task-specific fine-tuning, our zero-shot framework significantly outperforms state-of-the-art supervised methods across multiple benchmarks and demonstrates robust generalization to real-world settings.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14442
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
Zhang, Zixin
Chen, Kanghao
Wang, Hanqing
Zhang, Hongfei
Chen, Harold Haodong
Liao, Chenfei
Guo, Litao
Chen, Ying-Cong
Computer Vision and Pattern Recognition
Robotics
Affordance prediction, which identifies interaction regions on objects based on language instructions, is critical for embodied AI. Prevailing end-to-end models couple high-level reasoning and low-level grounding into a single monolithic pipeline and rely on training over annotated datasets, which leads to poor generalization on novel objects and unseen environments. In this paper, we move beyond this paradigm by proposing A4-Agent, a training-free agentic framework that decouples affordance prediction into a three-stage pipeline. Our framework coordinates specialized foundation models at test time: (1) a $\textbf{Dreamer}$ that employs generative models to visualize $\textit{how}$ an interaction would look; (2) a $\textbf{Thinker}$ that utilizes large vision-language models to decide $\textit{what}$ object part to interact with; and (3) a $\textbf{Spotter}$ that orchestrates vision foundation models to precisely locate $\textit{where}$ the interaction area is. By leveraging the complementary strengths of pre-trained models without any task-specific fine-tuning, our zero-shot framework significantly outperforms state-of-the-art supervised methods across multiple benchmarks and demonstrates robust generalization to real-world settings.
title A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2512.14442