ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Ying, Ling, Zhan, Shi, Haochen, Cheng, Jiayang, Yim, Yauwai, Song, Yangqiu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
by: Liang, Fangzhou, et al.
Published: (2025)
by: Liang, Fangzhou, et al.
Published: (2025)
NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding
by: Chan, Chunkit, et al.
Published: (2024)
by: Chan, Chunkit, et al.
Published: (2024)
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
by: Zheng, Tianshi, et al.
Published: (2024)
by: Zheng, Tianshi, et al.
Published: (2024)
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
by: Wang, Weiqi, et al.
Published: (2024)
by: Wang, Weiqi, et al.
Published: (2024)
AbsPyramid: Benchmarking the Abstraction Ability of Language Models with a Unified Entailment Graph
by: Wang, Zhaowei, et al.
Published: (2023)
by: Wang, Zhaowei, et al.
Published: (2023)
Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information
by: Yim, Yauwai, et al.
Published: (2024)
by: Yim, Yauwai, et al.
Published: (2024)
Persona Knowledge-Aligned Prompt Tuning Method for Online Debate
by: Chan, Chunkit, et al.
Published: (2024)
by: Chan, Chunkit, et al.
Published: (2024)
InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding
by: Jiayang, Cheng, et al.
Published: (2025)
by: Jiayang, Cheng, et al.
Published: (2025)
Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction
by: Deng, Zheye, et al.
Published: (2024)
by: Deng, Zheye, et al.
Published: (2024)
ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans
by: Sadana, Ananya, et al.
Published: (2025)
by: Sadana, Ananya, et al.
Published: (2025)
XToM: Exploring the Multilingual Theory of Mind for Large Language Models
by: Chan, Chunkit, et al.
Published: (2025)
by: Chan, Chunkit, et al.
Published: (2025)
PreAct: Prediction Enhances Agent's Planning Ability
by: Fu, Dayuan, et al.
Published: (2024)
by: Fu, Dayuan, et al.
Published: (2024)
LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning
by: Sun, Shibo, et al.
Published: (2025)
by: Sun, Shibo, et al.
Published: (2025)
OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset
by: Hu, Wenbin, et al.
Published: (2026)
by: Hu, Wenbin, et al.
Published: (2026)
CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense Reasoning
by: Wang, Weiqi, et al.
Published: (2024)
by: Wang, Weiqi, et al.
Published: (2024)
Anticipate & Act : Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments
by: Arora, Raghav, et al.
Published: (2025)
by: Arora, Raghav, et al.
Published: (2025)
LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
by: Jiayang, Cheng, et al.
Published: (2026)
by: Jiayang, Cheng, et al.
Published: (2026)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
by: Jing, Huihao, et al.
Published: (2026)
by: Jing, Huihao, et al.
Published: (2026)
Safety Compliance: Rethinking LLM Safety Reasoning through the Lens of Compliance
by: Hu, Wenbin, et al.
Published: (2025)
by: Hu, Wenbin, et al.
Published: (2025)
Monte Carlo Planning with Large Language Model for Text-Based Game Agents
by: Shi, Zijing, et al.
Published: (2025)
by: Shi, Zijing, et al.
Published: (2025)
EventGround: Narrative Reasoning by Grounding to Eventuality-centric Knowledge Graphs
by: Jiayang, Cheng, et al.
Published: (2024)
by: Jiayang, Cheng, et al.
Published: (2024)
IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Language Models in E-commerce
by: Ding, Wenxuan, et al.
Published: (2024)
by: Ding, Wenxuan, et al.
Published: (2024)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
by: Erdogan, Lutfi Eren, et al.
Published: (2025)
by: Erdogan, Lutfi Eren, et al.
Published: (2025)
PipeNet: Question Answering with Semantic Pruning over Knowledge Graphs
by: Su, Ying, et al.
Published: (2024)
by: Su, Ying, et al.
Published: (2024)
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
by: Xu, Baixuan, et al.
Published: (2025)
by: Xu, Baixuan, et al.
Published: (2025)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
by: Chen, Yi, et al.
Published: (2023)
by: Chen, Yi, et al.
Published: (2023)
GameTraversalBenchmark: Evaluating Planning Abilities Of Large Language Models Through Traversing 2D Game Maps
by: Nasir, Muhammad Umair, et al.
Published: (2024)
by: Nasir, Muhammad Umair, et al.
Published: (2024)
Distilling Instruction-following Abilities of Large Language Models with Task-aware Curriculum Planning
by: Yue, Yuanhao, et al.
Published: (2024)
by: Yue, Yuanhao, et al.
Published: (2024)
On the Ability of Transformers to Verify Plans
by: Sarrof, Yash, et al.
Published: (2026)
by: Sarrof, Yash, et al.
Published: (2026)
PARADISE: Evaluating Implicit Planning Skills of Language Models with Procedural Warnings and Tips Dataset
by: Uzunoglu, Arda, et al.
Published: (2024)
by: Uzunoglu, Arda, et al.
Published: (2024)
ChatGPT Evaluation on Sentence Level Relations: A Focus on Temporal, Causal, and Discourse Relations
by: Chan, Chunkit, et al.
Published: (2023)
by: Chan, Chunkit, et al.
Published: (2023)
Deliberate Planning in Language Models with Symbolic Representation
by: Xiong, Siheng, et al.
Published: (2025)
by: Xiong, Siheng, et al.
Published: (2025)
TravelPlanner: A Benchmark for Real-World Planning with Language Agents
by: Xie, Jian, et al.
Published: (2024)
by: Xie, Jian, et al.
Published: (2024)
PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
by: Fei, Zhaoye, et al.
Published: (2025)
by: Fei, Zhaoye, et al.
Published: (2025)
UrbanPlanBench: A Comprehensive Urban Planning Benchmark for Evaluating Large Language Models
by: Zheng, Yu, et al.
Published: (2025)
by: Zheng, Yu, et al.
Published: (2025)
Semformer: Transformer Language Models with Semantic Planning
by: Yin, Yongjing, et al.
Published: (2024)
by: Yin, Yongjing, et al.
Published: (2024)
EcomEdit: An Automated E-commerce Knowledge Editing Framework for Enhanced Product and Purchase Intention Understanding
by: Lau, Ching Ming Samuel, et al.
Published: (2024)
by: Lau, Ching Ming Samuel, et al.
Published: (2024)
PlanGPT: Enhancing Urban Planning with Tailored Language Model and Efficient Retrieval
by: Zhu, He, et al.
Published: (2024)
by: Zhu, He, et al.
Published: (2024)
Similar Items
-
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
by: Liang, Fangzhou, et al.
Published: (2025) -
NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding
by: Chan, Chunkit, et al.
Published: (2024) -
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
by: Zheng, Tianshi, et al.
Published: (2024) -
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
by: Wang, Weiqi, et al.
Published: (2024) -
AbsPyramid: Benchmarking the Abstraction Ability of Language Models with a Unified Entailment Graph
by: Wang, Zhaowei, et al.
Published: (2023)