Evaluating Long-Context Reasoning in LLM-Based WebAgents
Fuente:
arXiv
Salvato in:
| Autori principali: | Chung, Andy, Zhang, Yichi, Lin, Kaixiang, Rawal, Aditya, Gao, Qiaozi, Chai, Joyce |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
di: Gur, Izzeddin, et al.
Pubblicazione: (2023)
di: Gur, Izzeddin, et al.
Pubblicazione: (2023)
AgentFold: Long-Horizon Web Agents with Proactive Context Management
di: Ye, Rui, et al.
Pubblicazione: (2025)
di: Ye, Rui, et al.
Pubblicazione: (2025)
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
di: Zhang, Kehao, et al.
Pubblicazione: (2026)
di: Zhang, Kehao, et al.
Pubblicazione: (2026)
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)
Agentic Test-Time Scaling for WebAgents
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
di: Anupam, Sagnik, et al.
Pubblicazione: (2025)
di: Anupam, Sagnik, et al.
Pubblicazione: (2025)
Learning to Generate Formally Verifiable Step-by-Step Logic Reasoning via Structured Formal Intermediaries
di: Chen, Luoxin, et al.
Pubblicazione: (2026)
di: Chen, Luoxin, et al.
Pubblicazione: (2026)
Agentic Web: Weaving the Next Web with AI Agents
di: Yang, Yingxuan, et al.
Pubblicazione: (2025)
di: Yang, Yingxuan, et al.
Pubblicazione: (2025)
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
di: Yu, Hongli, et al.
Pubblicazione: (2025)
di: Yu, Hongli, et al.
Pubblicazione: (2025)
ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion
di: Fang, Zhenghan, et al.
Pubblicazione: (2025)
di: Fang, Zhenghan, et al.
Pubblicazione: (2025)
OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences
di: Wang, Kaixiang, et al.
Pubblicazione: (2026)
di: Wang, Kaixiang, et al.
Pubblicazione: (2026)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
Context, Reasoning, and Hierarchy: A Cost-Performance Study of Compound LLM Agent Design in an Adversarial POMDP
di: Bogdanov, Igor, et al.
Pubblicazione: (2026)
di: Bogdanov, Igor, et al.
Pubblicazione: (2026)
A Survey of WebAgents: Towards Next-Generation AI Agents for Web Automation with Large Foundation Models
di: Ning, Liangbo, et al.
Pubblicazione: (2025)
di: Ning, Liangbo, et al.
Pubblicazione: (2025)
Hierarchical Balance Packing: Towards Efficient Supervised Fine-tuning for Long-Context LLM
di: Yao, Yongqiang, et al.
Pubblicazione: (2025)
di: Yao, Yongqiang, et al.
Pubblicazione: (2025)
Agent-Omit: Adaptive Context Omission for Efficient LLM Agents
di: Ning, Yansong, et al.
Pubblicazione: (2026)
di: Ning, Yansong, et al.
Pubblicazione: (2026)
The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoning
di: Zhuang, Ren, et al.
Pubblicazione: (2026)
di: Zhuang, Ren, et al.
Pubblicazione: (2026)
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers
di: Ranganath, Aditya
Pubblicazione: (2026)
di: Ranganath, Aditya
Pubblicazione: (2026)
The PokeAgent Challenge: Competitive and Long-Context Learning at Scale
di: Karten, Seth, et al.
Pubblicazione: (2026)
di: Karten, Seth, et al.
Pubblicazione: (2026)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
di: Zhou, Ruijie, et al.
Pubblicazione: (2026)
di: Zhou, Ruijie, et al.
Pubblicazione: (2026)
Learning Long-Context Diffusion Policies via Past-Token Prediction
di: Torne, Marcel, et al.
Pubblicazione: (2025)
di: Torne, Marcel, et al.
Pubblicazione: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
di: Hadeliya, Tsimur, et al.
Pubblicazione: (2025)
di: Hadeliya, Tsimur, et al.
Pubblicazione: (2025)
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
di: Chaudhary, Siddharth, et al.
Pubblicazione: (2025)
di: Chaudhary, Siddharth, et al.
Pubblicazione: (2025)
MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning
di: Liu, Dong, et al.
Pubblicazione: (2026)
di: Liu, Dong, et al.
Pubblicazione: (2026)
FoldAct: Efficient and Stable Context Folding for Long-Horizon Search Agents
di: Shao, Jiaqi, et al.
Pubblicazione: (2025)
di: Shao, Jiaqi, et al.
Pubblicazione: (2025)
Graph of Agents: Principled Long Context Modeling by Emergent Multi-Agent Collaboration
di: Joo, Taejong, et al.
Pubblicazione: (2025)
di: Joo, Taejong, et al.
Pubblicazione: (2025)
DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning
di: Karaman, Batuhan K., et al.
Pubblicazione: (2026)
di: Karaman, Batuhan K., et al.
Pubblicazione: (2026)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
di: Ling Team, et al.
Pubblicazione: (2025)
di: Ling Team, et al.
Pubblicazione: (2025)
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
di: Li, Zhouyang, et al.
Pubblicazione: (2025)
di: Li, Zhouyang, et al.
Pubblicazione: (2025)
How to Train Your LLM Web Agent: A Statistical Diagnosis
di: Vattikonda, Dheeraj, et al.
Pubblicazione: (2025)
di: Vattikonda, Dheeraj, et al.
Pubblicazione: (2025)
Evaluating Very Long-Term Conversational Memory of LLM Agents
di: Maharana, Adyasha, et al.
Pubblicazione: (2024)
di: Maharana, Adyasha, et al.
Pubblicazione: (2024)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
di: Li, Miao, et al.
Pubblicazione: (2026)
di: Li, Miao, et al.
Pubblicazione: (2026)
Throttling Web Agents Using Reasoning Gates
di: Kumar, Abhinav, et al.
Pubblicazione: (2025)
di: Kumar, Abhinav, et al.
Pubblicazione: (2025)
On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems
di: Tang, Bohan, et al.
Pubblicazione: (2025)
di: Tang, Bohan, et al.
Pubblicazione: (2025)
PolicyLong: Towards On-Policy Context Extension
di: Jia, Junlong, et al.
Pubblicazione: (2026)
di: Jia, Junlong, et al.
Pubblicazione: (2026)
Evaluating the Ability of Explanations to Disambiguate Models in a Rashomon Set
di: Rawal, Kaivalya, et al.
Pubblicazione: (2026)
di: Rawal, Kaivalya, et al.
Pubblicazione: (2026)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
di: Li, Yan, et al.
Pubblicazione: (2026)
di: Li, Yan, et al.
Pubblicazione: (2026)
Boosting LLM Reasoning via Human-Inspired Reward Shaping
di: Lin, Wenze, et al.
Pubblicazione: (2026)
di: Lin, Wenze, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
di: Gur, Izzeddin, et al.
Pubblicazione: (2023) -
AgentFold: Long-Horizon Web Agents with Proactive Context Management
di: Ye, Rui, et al.
Pubblicazione: (2025) -
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
di: Zhang, Yichi, et al.
Pubblicazione: (2024) -
Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
di: Zhang, Kehao, et al.
Pubblicazione: (2026) -
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)