GroundAct: Can LLM Agents Ground Actions in Environmental States?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zixuan, Li, Dingming, Li, Hongxing, Miao, Yanrui, Chen, Shuo, Yan, Yuchen, Zhang, Wenqi, Shen, Yongliang, Lu, Weiming, Xiao, Jun, Zhuang, Yueting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910267727675392
author Wang, Zixuan
Li, Dingming
Li, Hongxing
Miao, Yanrui
Chen, Shuo
Yan, Yuchen
Zhang, Wenqi
Shen, Yongliang
Lu, Weiming
Xiao, Jun
Zhuang, Yueting
author_facet Wang, Zixuan
Li, Dingming
Li, Hongxing
Miao, Yanrui
Chen, Shuo
Yan, Yuchen
Zhang, Wenqi
Shen, Yongliang
Lu, Weiming
Xiao, Jun
Zhuang, Yueting
contents LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on environmental state that the instruction does not mention. We argue that this gap reflects a missing capability: action grounding, the ability to infer from structured environmental state whether an action is feasible, what prerequisites it lacks, and whether it exceeds individual capacity. We introduce GroundAct, a benchmark of 1,500 scenarios and 16,592 task instances in text-based interactive environments spanning 11 domains, with tasks organized into seven categories along a cognitive complexity hierarchy. Evaluating 15 LLMs (3B-671B), we find three diagnostic patterns: (i) attribute reasoning is weakly correlated with tool and coordination reasoning, producing distinct model profiles; (ii) complete environment graphs yield up to +27.6/-22.9% on tool use vs. implicit collaboration, separating search-bound from constraint-filtering bottlenecks; and (iii) supervised fine-tuning lifts Qwen2.5-3B from 0.6% to 76.3% on direct command but only 1.5% to 5.5% on implicit collaboration. These results establish action grounding as a multi-dimensional challenge irreducible to scaling.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05614
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GroundAct: Can LLM Agents Ground Actions in Environmental States?
Wang, Zixuan
Li, Dingming
Li, Hongxing
Miao, Yanrui
Chen, Shuo
Yan, Yuchen
Zhang, Wenqi
Shen, Yongliang
Lu, Weiming
Xiao, Jun
Zhuang, Yueting
Computation and Language
Artificial Intelligence
LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on environmental state that the instruction does not mention. We argue that this gap reflects a missing capability: action grounding, the ability to infer from structured environmental state whether an action is feasible, what prerequisites it lacks, and whether it exceeds individual capacity. We introduce GroundAct, a benchmark of 1,500 scenarios and 16,592 task instances in text-based interactive environments spanning 11 domains, with tasks organized into seven categories along a cognitive complexity hierarchy. Evaluating 15 LLMs (3B-671B), we find three diagnostic patterns: (i) attribute reasoning is weakly correlated with tool and coordination reasoning, producing distinct model profiles; (ii) complete environment graphs yield up to +27.6/-22.9% on tool use vs. implicit collaboration, separating search-bound from constraint-filtering bottlenecks; and (iii) supervised fine-tuning lifts Qwen2.5-3B from 0.6% to 76.3% on direct command but only 1.5% to 5.5% on implicit collaboration. These results establish action grounding as a multi-dimensional challenge irreducible to scaling.
title GroundAct: Can LLM Agents Ground Actions in Environmental States?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.05614