Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Aghzal, Mohamed, Stein, Gregory J., Yao, Ziyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey on Large Language Models for Automated Planning
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2025)
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2025)
Look Further Ahead: Testing the Limits of GPT-4 in Path Planning
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2024)
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2024)
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
von: Wang, Zehong, et al.
Veröffentlicht: (2026)
von: Wang, Zehong, et al.
Veröffentlicht: (2026)
Why Retrieval-Augmented Generation Fails: A Graph Perspective
von: Guo, Kai, et al.
Veröffentlicht: (2026)
von: Guo, Kai, et al.
Veröffentlicht: (2026)
Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2023)
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2023)
Evaluating Vision-Language Models as Evaluators in Path Planning
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2024)
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2024)
AI Planning Framework for LLM-Based Web Agents
von: Shahnovsky, Orit, et al.
Veröffentlicht: (2026)
von: Shahnovsky, Orit, et al.
Veröffentlicht: (2026)
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
von: Frank, Gregory N.
Veröffentlicht: (2026)
von: Frank, Gregory N.
Veröffentlicht: (2026)
HiPlan: Hierarchical Planning for LLM-Based Agents with Adaptive Global-Local Guidance
von: Li, Ziyue, et al.
Veröffentlicht: (2025)
von: Li, Ziyue, et al.
Veröffentlicht: (2025)
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
von: Yang, Ke, et al.
Veröffentlicht: (2024)
von: Yang, Ke, et al.
Veröffentlicht: (2024)
Why Chain of Thought Fails in Clinical Text Understanding
von: Wu, Jiageng, et al.
Veröffentlicht: (2025)
von: Wu, Jiageng, et al.
Veröffentlicht: (2025)
Automating the Generation of Prompts for LLM-based Action Choice in PDDL Planning
von: Stein, Katharina, et al.
Veröffentlicht: (2023)
von: Stein, Katharina, et al.
Veröffentlicht: (2023)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents
von: Zambrano, Alejandra, et al.
Veröffentlicht: (2026)
von: Zambrano, Alejandra, et al.
Veröffentlicht: (2026)
How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent
von: Jung, Sungwoo, et al.
Veröffentlicht: (2026)
von: Jung, Sungwoo, et al.
Veröffentlicht: (2026)
Why Do Multi-Agent LLM Systems Fail?
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
LLM Self-Explanations Fail Semantic Invariance
von: Szeider, Stefan
Veröffentlicht: (2026)
von: Szeider, Stefan
Veröffentlicht: (2026)
WebXSkill: Skill Learning for Autonomous Web Agents
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2026)
HoneyComb: A Flexible LLM-Based Agent System for Materials Science
von: Zhang, Huan, et al.
Veröffentlicht: (2024)
von: Zhang, Huan, et al.
Veröffentlicht: (2024)
WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk
von: Chen, Zichen, et al.
Veröffentlicht: (2025)
von: Chen, Zichen, et al.
Veröffentlicht: (2025)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
von: He, Hongliang, et al.
Veröffentlicht: (2024)
von: He, Hongliang, et al.
Veröffentlicht: (2024)
Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy
von: Pudasaini, Shushanta, et al.
Veröffentlicht: (2026)
von: Pudasaini, Shushanta, et al.
Veröffentlicht: (2026)
Causal Agent based on Large Language Model
von: Han, Kairong, et al.
Veröffentlicht: (2024)
von: Han, Kairong, et al.
Veröffentlicht: (2024)
Routine: A Structural Planning Framework for LLM Agent System in Enterprise
von: Zeng, Guancheng, et al.
Veröffentlicht: (2025)
von: Zeng, Guancheng, et al.
Veröffentlicht: (2025)
Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
NexusSum: Hierarchical LLM Agents for Long-Form Narrative Summarization
von: Kim, Hyuntak, et al.
Veröffentlicht: (2025)
von: Kim, Hyuntak, et al.
Veröffentlicht: (2025)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
von: Schwinn, Leo, et al.
Veröffentlicht: (2026)
von: Schwinn, Leo, et al.
Veröffentlicht: (2026)
HiMAP-Travel: Hierarchical Multi-Agent Planning for Long-Horizon Constrained Travel
von: Bui, The Viet, et al.
Veröffentlicht: (2026)
von: Bui, The Viet, et al.
Veröffentlicht: (2026)
Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation
von: Yu, Ye, et al.
Veröffentlicht: (2026)
von: Yu, Ye, et al.
Veröffentlicht: (2026)
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2025)
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2025)
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
von: Öncel, Fırat, et al.
Veröffentlicht: (2024)
von: Öncel, Fırat, et al.
Veröffentlicht: (2024)
Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
von: Chochlakis, Georgios, et al.
Veröffentlicht: (2024)
von: Chochlakis, Georgios, et al.
Veröffentlicht: (2024)
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
von: Yu, Tao, et al.
Veröffentlicht: (2025)
von: Yu, Tao, et al.
Veröffentlicht: (2025)
Automated Survey Collection with LLM-based Conversational Agents
von: Kaiyrbekov, Kurmanbek, et al.
Veröffentlicht: (2025)
von: Kaiyrbekov, Kurmanbek, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Survey on Large Language Models for Automated Planning
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2025) -
Look Further Ahead: Testing the Limits of GPT-4 in Path Planning
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2024) -
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
von: Wang, Zehong, et al.
Veröffentlicht: (2026) -
Why Retrieval-Augmented Generation Fails: A Graph Perspective
von: Guo, Kai, et al.
Veröffentlicht: (2026) -
Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2023)