GroundAct: Can LLM Agents Ground Actions in Environmental States?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910267727675392 |
|---|---|
| author | Wang, Zixuan Li, Dingming Li, Hongxing Miao, Yanrui Chen, Shuo Yan, Yuchen Zhang, Wenqi Shen, Yongliang Lu, Weiming Xiao, Jun Zhuang, Yueting |
| author_facet | Wang, Zixuan Li, Dingming Li, Hongxing Miao, Yanrui Chen, Shuo Yan, Yuchen Zhang, Wenqi Shen, Yongliang Lu, Weiming Xiao, Jun Zhuang, Yueting |
| contents | LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on environmental state that the instruction does not mention. We argue that this gap reflects a missing capability: action grounding, the ability to infer from structured environmental state whether an action is feasible, what prerequisites it lacks, and whether it exceeds individual capacity. We introduce GroundAct, a benchmark of 1,500 scenarios and 16,592 task instances in text-based interactive environments spanning 11 domains, with tasks organized into seven categories along a cognitive complexity hierarchy. Evaluating 15 LLMs (3B-671B), we find three diagnostic patterns: (i) attribute reasoning is weakly correlated with tool and coordination reasoning, producing distinct model profiles; (ii) complete environment graphs yield up to +27.6/-22.9% on tool use vs. implicit collaboration, separating search-bound from constraint-filtering bottlenecks; and (iii) supervised fine-tuning lifts Qwen2.5-3B from 0.6% to 76.3% on direct command but only 1.5% to 5.5% on implicit collaboration. These results establish action grounding as a multi-dimensional challenge irreducible to scaling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_05614 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GroundAct: Can LLM Agents Ground Actions in Environmental States? Wang, Zixuan Li, Dingming Li, Hongxing Miao, Yanrui Chen, Shuo Yan, Yuchen Zhang, Wenqi Shen, Yongliang Lu, Weiming Xiao, Jun Zhuang, Yueting Computation and Language Artificial Intelligence LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on environmental state that the instruction does not mention. We argue that this gap reflects a missing capability: action grounding, the ability to infer from structured environmental state whether an action is feasible, what prerequisites it lacks, and whether it exceeds individual capacity. We introduce GroundAct, a benchmark of 1,500 scenarios and 16,592 task instances in text-based interactive environments spanning 11 domains, with tasks organized into seven categories along a cognitive complexity hierarchy. Evaluating 15 LLMs (3B-671B), we find three diagnostic patterns: (i) attribute reasoning is weakly correlated with tool and coordination reasoning, producing distinct model profiles; (ii) complete environment graphs yield up to +27.6/-22.9% on tool use vs. implicit collaboration, separating search-bound from constraint-filtering bottlenecks; and (iii) supervised fine-tuning lifts Qwen2.5-3B from 0.6% to 76.3% on direct command but only 1.5% to 5.5% on implicit collaboration. These results establish action grounding as a multi-dimensional challenge irreducible to scaling. |
| title | GroundAct: Can LLM Agents Ground Actions in Environmental States? |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2508.05614 |