The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Xinyu Jessica, Bai, Haoyue, Sun, Yiyou, Wang, Haorui, Zhang, Shuibai, Hu, Wenjie, Schroder, Mya, Mutlu, Bilge, Song, Dawn, Nowak, Robert D |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How and Why LLMs Generalize: A Fine-Grained Analysis of LLM Reasoning from Cognitive Behaviors to Low-Level Patterns
por: Bai, Haoyue, et al.
Publicado: (2025)
por: Bai, Haoyue, et al.
Publicado: (2025)
Where's the liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated Content
por: Bai, Haoyue, et al.
Publicado: (2025)
por: Bai, Haoyue, et al.
Publicado: (2025)
LearnMate: Enhancing Online Education with LLM-Powered Personalized Learning Plans and Support
por: Wang, Xinyu Jessica, et al.
Publicado: (2025)
por: Wang, Xinyu Jessica, et al.
Publicado: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
por: Sun, Yiyou, et al.
Publicado: (2025)
por: Sun, Yiyou, et al.
Publicado: (2025)
LearnMate^2: Design and Evaluation of an LLM-powered Personalized and Adaptive Support System for Online Learning
por: Wang, Xinyu Jessica, et al.
Publicado: (2026)
por: Wang, Xinyu Jessica, et al.
Publicado: (2026)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
por: Sun, Yiyou, et al.
Publicado: (2025)
por: Sun, Yiyou, et al.
Publicado: (2025)
MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
por: Zhang, Weichen, et al.
Publicado: (2025)
por: Zhang, Weichen, et al.
Publicado: (2025)
Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
por: Lumer, Elias, et al.
Publicado: (2026)
por: Lumer, Elias, et al.
Publicado: (2026)
Where we left off Crowned Coil LLC
por: Brown, Mya
Publicado: (2026)
por: Brown, Mya
Publicado: (2026)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
por: Sun, Yiyou, et al.
Publicado: (2025)
por: Sun, Yiyou, et al.
Publicado: (2025)
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
por: Lee, Christine, et al.
Publicado: (2025)
por: Lee, Christine, et al.
Publicado: (2025)
Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
por: Lee, Yoonsang, et al.
Publicado: (2026)
por: Lee, Yoonsang, et al.
Publicado: (2026)
Toward Agentic AI: Task-Oriented Communication for Hierarchical Planning of Long-Horizon Tasks
por: Huang, Sin-Yu, et al.
Publicado: (2026)
por: Huang, Sin-Yu, et al.
Publicado: (2026)
U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning
por: Lee, Christine P, et al.
Publicado: (2026)
por: Lee, Christine P, et al.
Publicado: (2026)
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
por: Li, Ruiying, et al.
Publicado: (2026)
por: Li, Ruiying, et al.
Publicado: (2026)
Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought
por: Zhao, Jiachen, et al.
Publicado: (2025)
por: Zhao, Jiachen, et al.
Publicado: (2025)
AHA: Human-Assisted Out-of-Distribution Generalization and Detection
por: Bai, Haoyue, et al.
Publicado: (2024)
por: Bai, Haoyue, et al.
Publicado: (2024)
Toward Family-Robot Interactions: A Family-Centered Framework in HRI
por: Cagiltay, Bengisu, et al.
Publicado: (2024)
por: Cagiltay, Bengisu, et al.
Publicado: (2024)
Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
por: He, Shuo, et al.
Publicado: (2026)
por: He, Shuo, et al.
Publicado: (2026)
Sci-VLA: Agentic VLA Inference Plugin for Long-Horizon Tasks in Scientific Experiments
por: Pang, Yiwen, et al.
Publicado: (2026)
por: Pang, Yiwen, et al.
Publicado: (2026)
Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
por: Zhang, Yuxiang, et al.
Publicado: (2025)
por: Zhang, Yuxiang, et al.
Publicado: (2025)
Modular Jets for Supervised Pipelines: Diagnosing Mirage vs Identifiability
por: Sanyal, Suman
Publicado: (2025)
por: Sanyal, Suman
Publicado: (2025)
Crowdsourcing Task Traces for Service Robotics
por: Porfirio, David, et al.
Publicado: (2024)
por: Porfirio, David, et al.
Publicado: (2024)
Exploring the Use of Robots for Diary Studies
por: Xu, Michael F., et al.
Publicado: (2025)
por: Xu, Michael F., et al.
Publicado: (2025)
Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering
por: Zhu, Xinyu, et al.
Publicado: (2026)
por: Zhu, Xinyu, et al.
Publicado: (2026)
Diagnosing Hate Speech Classification: Where Do Humans and Machines Disagree, and Why?
por: Yang, Xilin
Publicado: (2024)
por: Yang, Xilin
Publicado: (2024)
Deep Active Learning in the Open World
por: Xie, Tian, et al.
Publicado: (2024)
por: Xie, Tian, et al.
Publicado: (2024)
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
por: Yin, Hao, et al.
Publicado: (2025)
por: Yin, Hao, et al.
Publicado: (2025)
MemWeaver: Weaving Hybrid Memories for Traceable Long-Horizon Agentic Reasoning
por: Ye, Juexiang, et al.
Publicado: (2026)
por: Ye, Juexiang, et al.
Publicado: (2026)
MAP: Multi-user Personalization with Collaborative LLM-powered Agents
por: Lee, Christine, et al.
Publicado: (2025)
por: Lee, Christine, et al.
Publicado: (2025)
Sprout: Designing Expressivity for Robots Using Fiber-Embedded Actuator
por: Koike, Amy, et al.
Publicado: (2024)
por: Koike, Amy, et al.
Publicado: (2024)
Factors that Affect Personalization of Robots for Older Adults
por: Stegner, Laura, et al.
Publicado: (2024)
por: Stegner, Laura, et al.
Publicado: (2024)
Tangible Scenography as a Holistic Design Method for Human-Robot Interaction
por: Koike, Amy, et al.
Publicado: (2024)
por: Koike, Amy, et al.
Publicado: (2024)
Towards Robust Out-of-Distribution Generalization: Data Augmentation and Neural Architecture Search Approaches
por: Bai, Haoyue
Publicado: (2024)
por: Bai, Haoyue
Publicado: (2024)
Strategy Executability in Mathematical Reasoning: Leveraging Human-Model Differences for Effective Guidance
por: Liang, Weida, et al.
Publicado: (2026)
por: Liang, Weida, et al.
Publicado: (2026)
Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
por: Wu, Haoze, et al.
Publicado: (2025)
por: Wu, Haoze, et al.
Publicado: (2025)
Agentic Feature Augmentation: Unifying Selection and Generation with Teaming, Planning, and Memories
por: Gong, Nanxu, et al.
Publicado: (2025)
por: Gong, Nanxu, et al.
Publicado: (2025)
Laser: Governing Long-Horizon Agentic Search via Structured Protocol and Context Register
por: Wang, Shuting, et al.
Publicado: (2025)
por: Wang, Shuting, et al.
Publicado: (2025)
Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization
por: Chen, Qianben, et al.
Publicado: (2026)
por: Chen, Qianben, et al.
Publicado: (2026)
LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning
por: Pushkin, Denys, et al.
Publicado: (2026)
por: Pushkin, Denys, et al.
Publicado: (2026)
Ejemplares similares
-
How and Why LLMs Generalize: A Fine-Grained Analysis of LLM Reasoning from Cognitive Behaviors to Low-Level Patterns
por: Bai, Haoyue, et al.
Publicado: (2025) -
Where's the liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated Content
por: Bai, Haoyue, et al.
Publicado: (2025) -
LearnMate: Enhancing Online Education with LLM-Powered Personalized Learning Plans and Support
por: Wang, Xinyu Jessica, et al.
Publicado: (2025) -
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
por: Sun, Yiyou, et al.
Publicado: (2025) -
LearnMate^2: Design and Evaluation of an LLM-powered Personalized and Adaptive Support System for Online Learning
por: Wang, Xinyu Jessica, et al.
Publicado: (2026)