The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Sinha, Akshit, Arun, Arvindh, Goel, Shashwat, Staab, Steffen, Geiping, Jonas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FutureSim: Replaying World Events to Evaluate Adaptive Agents
di: Goel, Shashwat, et al.
Pubblicazione: (2026)
di: Goel, Shashwat, et al.
Pubblicazione: (2026)
A Cognac Shot To Forget Bad Memories: Corrective Unlearning for Graph Neural Networks
di: Kolipaka, Varshita, et al.
Pubblicazione: (2024)
di: Kolipaka, Varshita, et al.
Pubblicazione: (2024)
The Illusion of Procedural Reasoning: Measuring Long-Horizon FSM Execution in LLMs
di: Samiei, Mahdi, et al.
Pubblicazione: (2025)
di: Samiei, Mahdi, et al.
Pubblicazione: (2025)
Pitfalls in Evaluating Language Model Forecasters
di: Paleka, Daniel, et al.
Pubblicazione: (2025)
di: Paleka, Daniel, et al.
Pubblicazione: (2025)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
di: Chandak, Nikhil, et al.
Pubblicazione: (2025)
di: Chandak, Nikhil, et al.
Pubblicazione: (2025)
SEMMA: A Semantic Aware Knowledge Graph Foundation Model
di: Arun, Arvindh, et al.
Pubblicazione: (2025)
di: Arun, Arvindh, et al.
Pubblicazione: (2025)
The Diminishing Returns of Early-Exit Decoding in Modern LLMs
di: Wei, Rui, et al.
Pubblicazione: (2026)
di: Wei, Rui, et al.
Pubblicazione: (2026)
Intrinsic Credit Assignment for Long Horizon Interaction
di: Auzina, Ilze Amanda, et al.
Pubblicazione: (2026)
di: Auzina, Ilze Amanda, et al.
Pubblicazione: (2026)
On the Diminishing Returns of Width for Continual Learning
di: Guha, Etash, et al.
Pubblicazione: (2024)
di: Guha, Etash, et al.
Pubblicazione: (2024)
LLMs4Life: Large Language Models for Ontology Learning in Life Sciences
di: Fathallah, Nadeen, et al.
Pubblicazione: (2024)
di: Fathallah, Nadeen, et al.
Pubblicazione: (2024)
Self-Consistency Is Losing Its Edge: Diminishing Returns and Rising Costs in Modern LLMs
di: Loo, Chiyan
Pubblicazione: (2025)
di: Loo, Chiyan
Pubblicazione: (2025)
AccessGuru: Leveraging LLMs to Detect and Correct Web Accessibility Violations in HTML Code
di: Fathallah, Nadeen, et al.
Pubblicazione: (2025)
di: Fathallah, Nadeen, et al.
Pubblicazione: (2025)
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging
di: Su, Guinan, et al.
Pubblicazione: (2025)
di: Su, Guinan, et al.
Pubblicazione: (2025)
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
di: Hu, Xavier, et al.
Pubblicazione: (2026)
di: Hu, Xavier, et al.
Pubblicazione: (2026)
From Tokens to Lattices: Emergent Lattice Structures in Language Models
di: Xiong, Bo, et al.
Pubblicazione: (2025)
di: Xiong, Bo, et al.
Pubblicazione: (2025)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
di: Sharma, Agniv, et al.
Pubblicazione: (2024)
di: Sharma, Agniv, et al.
Pubblicazione: (2024)
Great Models Think Alike and this Undermines AI Oversight
di: Goel, Shashwat, et al.
Pubblicazione: (2025)
di: Goel, Shashwat, et al.
Pubblicazione: (2025)
Empowering the Deaf and Hard of Hearing Community: Enhancing Video Captions Using Large Language Models
di: Fathallah, Nadeen, et al.
Pubblicazione: (2024)
di: Fathallah, Nadeen, et al.
Pubblicazione: (2024)
CAFIN: Centrality Aware Fairness inducing IN-processing for Unsupervised Representation Learning on Graphs
di: Arun, Arvindh, et al.
Pubblicazione: (2023)
di: Arun, Arvindh, et al.
Pubblicazione: (2023)
Intrinsic Stability Limits of Autoregressive Reasoning: Structural Consequences for Long-Horizon Execution
di: Liao, Hsien-Jyh
Pubblicazione: (2026)
di: Liao, Hsien-Jyh
Pubblicazione: (2026)
The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
di: Sinha, Debu
Pubblicazione: (2025)
di: Sinha, Debu
Pubblicazione: (2025)
PARC: An Autonomous Self-Reflective Coding Agent for Robust Execution of Long-Horizon Tasks
di: Orimo, Yuki, et al.
Pubblicazione: (2025)
di: Orimo, Yuki, et al.
Pubblicazione: (2025)
SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
di: Li, Jialiang, et al.
Pubblicazione: (2025)
di: Li, Jialiang, et al.
Pubblicazione: (2025)
F -- A Model of Events based on the Foundational Ontology DOLCE+DnS Ultralite
di: Scherp, Ansgar, et al.
Pubblicazione: (2024)
di: Scherp, Ansgar, et al.
Pubblicazione: (2024)
Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation
di: Deng, Zehao, et al.
Pubblicazione: (2025)
di: Deng, Zehao, et al.
Pubblicazione: (2025)
When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution
di: Zhu, Zilin, et al.
Pubblicazione: (2026)
di: Zhu, Zilin, et al.
Pubblicazione: (2026)
Cooperative Evolutionary Pressure and Diminishing Returns Might Explain the Fermi Paradox: On What Super-AIs Are Like
di: Vallstrom, Daniel
Pubblicazione: (2024)
di: Vallstrom, Daniel
Pubblicazione: (2024)
Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning
di: Hwang, Jaebak, et al.
Pubblicazione: (2025)
di: Hwang, Jaebak, et al.
Pubblicazione: (2025)
Can LLMs Introspect? A Reality Check
di: Singh, Shashwat, et al.
Pubblicazione: (2026)
di: Singh, Shashwat, et al.
Pubblicazione: (2026)
Expanding Expressivity in Transformer Models with MöbiusAttention
di: Halacheva, Anna-Maria, et al.
Pubblicazione: (2024)
di: Halacheva, Anna-Maria, et al.
Pubblicazione: (2024)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
di: Elhassan, Fay, et al.
Pubblicazione: (2025)
di: Elhassan, Fay, et al.
Pubblicazione: (2025)
Models That Know How Evaluations Are Designed Score Safer
di: Deckenbach, Katharina, et al.
Pubblicazione: (2026)
di: Deckenbach, Katharina, et al.
Pubblicazione: (2026)
Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments
di: Goel, Hardik
Pubblicazione: (2026)
di: Goel, Hardik
Pubblicazione: (2026)
Performance of LLMs on Stochastic Modeling Operations Research Problems: From Theory to Practice
di: Kumar, Akshit, et al.
Pubblicazione: (2025)
di: Kumar, Akshit, et al.
Pubblicazione: (2025)
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
di: Ezra, Elon, et al.
Pubblicazione: (2025)
di: Ezra, Elon, et al.
Pubblicazione: (2025)
Technical Debt in In-Context Learning: Diminishing Efficiency in Long Context
di: Joo, Taejong, et al.
Pubblicazione: (2025)
di: Joo, Taejong, et al.
Pubblicazione: (2025)
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation
di: Sinha, Shiven, et al.
Pubblicazione: (2025)
di: Sinha, Shiven, et al.
Pubblicazione: (2025)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
di: Wang, Keyu, et al.
Pubblicazione: (2025)
di: Wang, Keyu, et al.
Pubblicazione: (2025)
SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio
di: Pandey, Satwik, et al.
Pubblicazione: (2026)
di: Pandey, Satwik, et al.
Pubblicazione: (2026)
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
di: Li, Junlong, et al.
Pubblicazione: (2025)
di: Li, Junlong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FutureSim: Replaying World Events to Evaluate Adaptive Agents
di: Goel, Shashwat, et al.
Pubblicazione: (2026) -
A Cognac Shot To Forget Bad Memories: Corrective Unlearning for Graph Neural Networks
di: Kolipaka, Varshita, et al.
Pubblicazione: (2024) -
The Illusion of Procedural Reasoning: Measuring Long-Horizon FSM Execution in LLMs
di: Samiei, Mahdi, et al.
Pubblicazione: (2025) -
Pitfalls in Evaluating Language Model Forecasters
di: Paleka, Daniel, et al.
Pubblicazione: (2025) -
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
di: Chandak, Nikhil, et al.
Pubblicazione: (2025)