Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks
Fuente:
arXiv
Saved in:
| Main Author: | Lee, Wilson Y. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
by: Lourie, Nicholas, et al.
Published: (2025)
by: Lourie, Nicholas, et al.
Published: (2025)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
by: Jang, Lawrence Keunho, et al.
Published: (2026)
by: Jang, Lawrence Keunho, et al.
Published: (2026)
LHAW: Controllable Underspecification for Long-Horizon Tasks
by: Pu, George, et al.
Published: (2026)
by: Pu, George, et al.
Published: (2026)
Mem-T: Densifying Rewards for Long-Horizon Memory Agents
by: Yue, Yanwei, et al.
Published: (2026)
by: Yue, Yanwei, et al.
Published: (2026)
Scaling Long-Horizon LLM Agent via Context-Folding
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
by: Wang, Jiawei, et al.
Published: (2025)
by: Wang, Jiawei, et al.
Published: (2025)
Intrinsic Language-Guided Exploration for Complex Long-Horizon Robotic Manipulation Tasks
by: Triantafyllidis, Eleftherios, et al.
Published: (2023)
by: Triantafyllidis, Eleftherios, et al.
Published: (2023)
AgentFold: Long-Horizon Web Agents with Proactive Context Management
by: Ye, Rui, et al.
Published: (2025)
by: Ye, Rui, et al.
Published: (2025)
Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
by: Wu, Xixi, et al.
Published: (2026)
by: Wu, Xixi, et al.
Published: (2026)
Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory
by: Wang, Zhenting, et al.
Published: (2026)
by: Wang, Zhenting, et al.
Published: (2026)
How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation
by: Lee, Wilson Y.
Published: (2026)
by: Lee, Wilson Y.
Published: (2026)
Recursive Models for Long-Horizon Reasoning
by: Yang, Chenxiao, et al.
Published: (2026)
by: Yang, Chenxiao, et al.
Published: (2026)
Failure Modes of LLMs for Causal Reasoning on Narratives
by: Yamin, Khurram, et al.
Published: (2024)
by: Yamin, Khurram, et al.
Published: (2024)
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
by: Yeo, Woongyeng, et al.
Published: (2026)
by: Yeo, Woongyeng, et al.
Published: (2026)
Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention
by: Vasudev, Rakshith, et al.
Published: (2026)
by: Vasudev, Rakshith, et al.
Published: (2026)
Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities
by: Spiegelhalter, Urs, et al.
Published: (2025)
by: Spiegelhalter, Urs, et al.
Published: (2025)
PRInTS: Reward Modeling for Long-Horizon Information Seeking
by: Lee, Jaewoo, et al.
Published: (2025)
by: Lee, Jaewoo, et al.
Published: (2025)
Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles
by: Shah, Kulin, et al.
Published: (2024)
by: Shah, Kulin, et al.
Published: (2024)
The Mystery of the Pathological Path-star Task for Language Models
by: Frydenlund, Arvid
Published: (2024)
by: Frydenlund, Arvid
Published: (2024)
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
by: Choi, Hyeonje, et al.
Published: (2026)
by: Choi, Hyeonje, et al.
Published: (2026)
Tool Preferences in Agentic LLMs are Unreliable
by: Faghih, Kazem, et al.
Published: (2025)
by: Faghih, Kazem, et al.
Published: (2025)
CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models
by: Tu, Ruibo, et al.
Published: (2024)
by: Tu, Ruibo, et al.
Published: (2024)
Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
by: Kiruluta, Andrew, et al.
Published: (2025)
by: Kiruluta, Andrew, et al.
Published: (2025)
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
by: Wu, Mingqi, et al.
Published: (2025)
by: Wu, Mingqi, et al.
Published: (2025)
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
by: Wang, Zehong, et al.
Published: (2026)
by: Wang, Zehong, et al.
Published: (2026)
Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning
by: Wu, Jinyang, et al.
Published: (2026)
by: Wu, Jinyang, et al.
Published: (2026)
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
by: Ge, Tao, et al.
Published: (2026)
by: Ge, Tao, et al.
Published: (2026)
The NLP Task Effectiveness of Long-Range Transformers
by: Qin, Guanghui, et al.
Published: (2022)
by: Qin, Guanghui, et al.
Published: (2022)
How does Multi-Task Training Affect Transformer In-Context Capabilities? Investigations with Function Classes
by: Bhasin, Harmon, et al.
Published: (2024)
by: Bhasin, Harmon, et al.
Published: (2024)
Investigating Symbolic Capabilities of Large Language Models
by: Dave, Neisarg, et al.
Published: (2024)
by: Dave, Neisarg, et al.
Published: (2024)
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
by: Li, Chenliang, et al.
Published: (2025)
by: Li, Chenliang, et al.
Published: (2025)
Enhancing LLM Agent Safety via Causal Influence Prompting
by: Hahm, Dongyoon, et al.
Published: (2025)
by: Hahm, Dongyoon, et al.
Published: (2025)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
by: Braun, Joschka
Published: (2026)
by: Braun, Joschka
Published: (2026)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
by: Li, Guanlin, et al.
Published: (2025)
by: Li, Guanlin, et al.
Published: (2025)
The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
by: Hadeliya, Tsimur, et al.
Published: (2025)
by: Hadeliya, Tsimur, et al.
Published: (2025)
Diversity Enhances an LLM's Performance in RAG and Long-context Task
by: Wang, Zhichao, et al.
Published: (2025)
by: Wang, Zhichao, et al.
Published: (2025)
MuPlon: Multi-Path Causal Optimization for Claim Verification through Controlling Confounding
by: Guo, Hanghui, et al.
Published: (2025)
by: Guo, Hanghui, et al.
Published: (2025)
Similar Items
-
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
by: Lourie, Nicholas, et al.
Published: (2025) -
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
by: Jang, Lawrence Keunho, et al.
Published: (2026) -
LHAW: Controllable Underspecification for Long-Horizon Tasks
by: Pu, George, et al.
Published: (2026) -
Mem-T: Densifying Rewards for Long-Horizon Memory Agents
by: Yue, Yanwei, et al.
Published: (2026) -
Scaling Long-Horizon LLM Agent via Context-Folding
by: Sun, Weiwei, et al.
Published: (2025)