ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Pei-An, Liang, Yong-Ching, Yeh, Jia-Fong, Su, Hung-Ting, Chen, Yi-Ting, Sun, Min, Hsu, Winston
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913047036035072
author Chen, Pei-An
Liang, Yong-Ching
Yeh, Jia-Fong
Su, Hung-Ting
Chen, Yi-Ting
Sun, Min
Hsu, Winston
author_facet Chen, Pei-An
Liang, Yong-Ching
Yeh, Jia-Fong
Su, Hung-Ting
Chen, Yi-Ting
Sun, Min
Hsu, Winston
contents Intelligent embodied agents should not simply follow instructions, as real-world environments often involve unexpected conditions and exceptions. However, existing methods usually focus on directly executing instructions, without considering whether the target objects can actually be manipulated, meaning they fail to assess available affordances. To address this limitation, we introduce DynAfford, a benchmark that evaluates embodied agents in dynamic environments where object affordances may change over time and are not specified in the instruction. DynAfford requires agents to perceive object states, infer implicit preconditions, and adapt their actions accordingly. To enable this capability, we introduce ADAPT, a plug-and-play module that augments existing planners with explicit affordance reasoning. Experiments demonstrate that incorporating ADAPT significantly improves robustness and task success across both seen and unseen environments. We also show that a domain-adapted, LoRA-finetuned vision-language model used as the affordance inference backend outperforms a commercial LLM (GPT-4o), highlighting the importance of task-aligned affordance grounding.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14902
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
Chen, Pei-An
Liang, Yong-Ching
Yeh, Jia-Fong
Su, Hung-Ting
Chen, Yi-Ting
Sun, Min
Hsu, Winston
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Robotics
Intelligent embodied agents should not simply follow instructions, as real-world environments often involve unexpected conditions and exceptions. However, existing methods usually focus on directly executing instructions, without considering whether the target objects can actually be manipulated, meaning they fail to assess available affordances. To address this limitation, we introduce DynAfford, a benchmark that evaluates embodied agents in dynamic environments where object affordances may change over time and are not specified in the instruction. DynAfford requires agents to perceive object states, infer implicit preconditions, and adapt their actions accordingly. To enable this capability, we introduce ADAPT, a plug-and-play module that augments existing planners with explicit affordance reasoning. Experiments demonstrate that incorporating ADAPT significantly improves robustness and task success across both seen and unseen environments. We also show that a domain-adapted, LoRA-finetuned vision-language model used as the affordance inference backend outperforms a commercial LLM (GPT-4o), highlighting the importance of task-aligned affordance grounding.
title ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2604.14902