Haste Makes Waste: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Zirui, Liu, Xiao, Li, Jiayi, Kong, Lingpeng, Feng, Yansong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ProTrix: Building Models for Planning and Reasoning over Tables with Sentence Context
di: Wu, Zirui, et al.
Pubblicazione: (2024)
di: Wu, Zirui, et al.
Pubblicazione: (2024)
Reasoning Does Not Necessarily Improve Role-Playing Ability
di: Feng, Xiachong, et al.
Pubblicazione: (2025)
di: Feng, Xiachong, et al.
Pubblicazione: (2025)
Automated Annotation of Evolving Corpora for Augmenting Longitudinal Network Data: A Framework Integrating Large Language Models and Expert Knowledge
di: Liu, Xiao, et al.
Pubblicazione: (2025)
di: Liu, Xiao, et al.
Pubblicazione: (2025)
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
di: Liu, Xiao, et al.
Pubblicazione: (2025)
di: Liu, Xiao, et al.
Pubblicazione: (2025)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
di: Liu, Xiao, et al.
Pubblicazione: (2024)
di: Liu, Xiao, et al.
Pubblicazione: (2024)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
di: Jiang, Botian, et al.
Pubblicazione: (2024)
di: Jiang, Botian, et al.
Pubblicazione: (2024)
GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
di: Li, Qintong, et al.
Pubblicazione: (2024)
di: Li, Qintong, et al.
Pubblicazione: (2024)
DreamOn: Diffusion Language Models For Code Infilling Beyond Fixed-size Canvas
di: Wu, Zirui, et al.
Pubblicazione: (2026)
di: Wu, Zirui, et al.
Pubblicazione: (2026)
DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)
D$^2$Plan: Dual-Agent Dynamic Global Planning for Complex Retrieval-Augmented Reasoning
di: Luo, Kangcheng, et al.
Pubblicazione: (2026)
di: Luo, Kangcheng, et al.
Pubblicazione: (2026)
DiNeR: a Large Realistic Dataset for Evaluating Compositional Generalization
di: Hu, Chengang, et al.
Pubblicazione: (2024)
di: Hu, Chengang, et al.
Pubblicazione: (2024)
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks
di: Li, Qintong, et al.
Pubblicazione: (2023)
di: Li, Qintong, et al.
Pubblicazione: (2023)
FACTTRACK: Time-Aware World State Tracking in Story Outlines
di: Lyu, Zhiheng, et al.
Pubblicazione: (2024)
di: Lyu, Zhiheng, et al.
Pubblicazione: (2024)
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
di: Huang, Jen-tse, et al.
Pubblicazione: (2024)
di: Huang, Jen-tse, et al.
Pubblicazione: (2024)
TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation
di: Zhang, Yikai, et al.
Pubblicazione: (2024)
di: Zhang, Yikai, et al.
Pubblicazione: (2024)
Dream 7B: Diffusion Large Language Models
di: Ye, Jiacheng, et al.
Pubblicazione: (2025)
di: Ye, Jiacheng, et al.
Pubblicazione: (2025)
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation
di: Luo, Kangcheng, et al.
Pubblicazione: (2025)
di: Luo, Kangcheng, et al.
Pubblicazione: (2025)
Scaling Reasoning without Attention
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)
Proxy Compression for Language Modeling
di: Zheng, Lin, et al.
Pubblicazione: (2026)
di: Zheng, Lin, et al.
Pubblicazione: (2026)
Non-myopic Generation of Language Models for Reasoning and Planning
di: Ma, Chang, et al.
Pubblicazione: (2024)
di: Ma, Chang, et al.
Pubblicazione: (2024)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
di: Chang, Jiayi, et al.
Pubblicazione: (2025)
di: Chang, Jiayi, et al.
Pubblicazione: (2025)
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
di: Wang, Haochuan, et al.
Pubblicazione: (2024)
di: Wang, Haochuan, et al.
Pubblicazione: (2024)
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
di: Lin, Jiuheng, et al.
Pubblicazione: (2025)
di: Lin, Jiuheng, et al.
Pubblicazione: (2025)
LaTeX Compilation: Challenges in the Era of LLMs
di: Liu, Tianyou, et al.
Pubblicazione: (2026)
di: Liu, Tianyou, et al.
Pubblicazione: (2026)
Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling
di: Wang, Xinglin, et al.
Pubblicazione: (2026)
di: Wang, Xinglin, et al.
Pubblicazione: (2026)
Can Perplexity Reflect Large Language Model's Ability in Long Text Understanding?
di: Hu, Yutong, et al.
Pubblicazione: (2024)
di: Hu, Yutong, et al.
Pubblicazione: (2024)
How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models
di: Jiang, Jiyue, et al.
Pubblicazione: (2024)
di: Jiang, Jiyue, et al.
Pubblicazione: (2024)
ELLA: Empowering LLMs for Interpretable, Accurate and Informative Legal Advice
di: Hu, Yutong, et al.
Pubblicazione: (2024)
di: Hu, Yutong, et al.
Pubblicazione: (2024)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
di: Kuang, Jiayi, et al.
Pubblicazione: (2025)
di: Kuang, Jiayi, et al.
Pubblicazione: (2025)
Forewarned is Forearmed: Leveraging LLMs for Data Synthesis through Failure-Inducing Exploration
di: Li, Qintong, et al.
Pubblicazione: (2024)
di: Li, Qintong, et al.
Pubblicazione: (2024)
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
di: Cui, Hao, et al.
Pubblicazione: (2025)
di: Cui, Hao, et al.
Pubblicazione: (2025)
P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
di: Zhang, Yidan, et al.
Pubblicazione: (2024)
di: Zhang, Yidan, et al.
Pubblicazione: (2024)
Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
di: Ye, Jiacheng, et al.
Pubblicazione: (2024)
di: Ye, Jiacheng, et al.
Pubblicazione: (2024)
Why Does the Effective Context Length of LLMs Fall Short?
di: An, Chenxin, et al.
Pubblicazione: (2024)
di: An, Chenxin, et al.
Pubblicazione: (2024)
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
di: Fei, Zhaoye, et al.
Pubblicazione: (2025)
di: Fei, Zhaoye, et al.
Pubblicazione: (2025)
Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling
di: Prange, Jakob, et al.
Pubblicazione: (2021)
di: Prange, Jakob, et al.
Pubblicazione: (2021)
Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively
di: Gu, Jiawei, et al.
Pubblicazione: (2025)
di: Gu, Jiawei, et al.
Pubblicazione: (2025)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
di: Liu, Yijun, et al.
Pubblicazione: (2024)
di: Liu, Yijun, et al.
Pubblicazione: (2024)
CASA: Causality-driven Argument Sufficiency Assessment
di: Liu, Xiao, et al.
Pubblicazione: (2024)
di: Liu, Xiao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ProTrix: Building Models for Planning and Reasoning over Tables with Sentence Context
di: Wu, Zirui, et al.
Pubblicazione: (2024) -
Reasoning Does Not Necessarily Improve Role-Playing Ability
di: Feng, Xiachong, et al.
Pubblicazione: (2025) -
Automated Annotation of Evolving Corpora for Augmenting Longitudinal Network Data: A Framework Integrating Large Language Models and Expert Knowledge
di: Liu, Xiao, et al.
Pubblicazione: (2025) -
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
di: Liu, Xiao, et al.
Pubblicazione: (2025) -
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
di: Liu, Xiao, et al.
Pubblicazione: (2024)