How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent
Fuente:
arXiv
Salvato in:
| Autori principali: | Jung, Sungwoo, Son, Seonil |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Much Can RAG Help the Reasoning of LLM?
di: Liu, Jingyu, et al.
Pubblicazione: (2024)
di: Liu, Jingyu, et al.
Pubblicazione: (2024)
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
di: Islam, Saad Obaid ul, et al.
Pubblicazione: (2025)
di: Islam, Saad Obaid ul, et al.
Pubblicazione: (2025)
Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective
di: Aghzal, Mohamed, et al.
Pubblicazione: (2026)
di: Aghzal, Mohamed, et al.
Pubblicazione: (2026)
Code as Agent Harness
di: Ning, Xuying, et al.
Pubblicazione: (2026)
di: Ning, Xuying, et al.
Pubblicazione: (2026)
It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers
di: Cho, Yong-eun
Pubblicazione: (2026)
di: Cho, Yong-eun
Pubblicazione: (2026)
Natural-Language Agent Harnesses
di: Pan, Linyue, et al.
Pubblicazione: (2026)
di: Pan, Linyue, et al.
Pubblicazione: (2026)
How Much Can We Forget about Data Contamination?
di: Bordt, Sebastian, et al.
Pubblicazione: (2024)
di: Bordt, Sebastian, et al.
Pubblicazione: (2024)
Do Large Language Models Know How Much They Know?
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Task--Specificity Score: Measuring How Much Instructions Really Matter for Supervision
di: Kadasi, Pritam, et al.
Pubblicazione: (2026)
di: Kadasi, Pritam, et al.
Pubblicazione: (2026)
Arena-Lite: Efficient and Reliable Large Language Model Evaluation via Tournament-Based Direct Comparisons
di: Son, Seonil, et al.
Pubblicazione: (2024)
di: Son, Seonil, et al.
Pubblicazione: (2024)
ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild
di: Yang, Bufang, et al.
Pubblicazione: (2025)
di: Yang, Bufang, et al.
Pubblicazione: (2025)
AI Planning Framework for LLM-Based Web Agents
di: Shahnovsky, Orit, et al.
Pubblicazione: (2026)
di: Shahnovsky, Orit, et al.
Pubblicazione: (2026)
Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation
di: Yu, Ye, et al.
Pubblicazione: (2026)
di: Yu, Ye, et al.
Pubblicazione: (2026)
UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind
di: Qian, Cheng, et al.
Pubblicazione: (2026)
di: Qian, Cheng, et al.
Pubblicazione: (2026)
s3: You Don't Need That Much Data to Train a Search Agent via RL
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
di: Gioacchini, Luca, et al.
Pubblicazione: (2024)
di: Gioacchini, Luca, et al.
Pubblicazione: (2024)
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
di: Gonzalez-Pumariega, Gonzalo, et al.
Pubblicazione: (2025)
di: Gonzalez-Pumariega, Gonzalo, et al.
Pubblicazione: (2025)
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
di: Jung, Jimin, et al.
Pubblicazione: (2026)
di: Jung, Jimin, et al.
Pubblicazione: (2026)
From Biased Chatbots to Biased Agents: Examining Role Assignment Effects on LLM Agent Robustness
di: Cao, Linbo, et al.
Pubblicazione: (2026)
di: Cao, Linbo, et al.
Pubblicazione: (2026)
GroundAct: Can LLM Agents Ground Actions in Environmental States?
di: Wang, Zixuan, et al.
Pubblicazione: (2025)
di: Wang, Zixuan, et al.
Pubblicazione: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
HiPlan: Hierarchical Planning for LLM-Based Agents with Adaptive Global-Local Guidance
di: Li, Ziyue, et al.
Pubblicazione: (2025)
di: Li, Ziyue, et al.
Pubblicazione: (2025)
How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting
di: Seegmiller, Parker, et al.
Pubblicazione: (2026)
di: Seegmiller, Parker, et al.
Pubblicazione: (2026)
VeRO: An Evaluation Harness for Agents to Optimize Agents
di: Ursekar, Varun, et al.
Pubblicazione: (2026)
di: Ursekar, Varun, et al.
Pubblicazione: (2026)
Stance Detection with Collaborative Role-Infused LLM-Based Agents
di: Lan, Xiaochong, et al.
Pubblicazione: (2023)
di: Lan, Xiaochong, et al.
Pubblicazione: (2023)
BertaQA: How Much Do Language Models Know About Local Culture?
di: Etxaniz, Julen, et al.
Pubblicazione: (2024)
di: Etxaniz, Julen, et al.
Pubblicazione: (2024)
FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration
di: Jung, Dongwon, et al.
Pubblicazione: (2025)
di: Jung, Dongwon, et al.
Pubblicazione: (2025)
Routine: A Structural Planning Framework for LLM Agent System in Enterprise
di: Zeng, Guancheng, et al.
Pubblicazione: (2025)
di: Zeng, Guancheng, et al.
Pubblicazione: (2025)
Diagnosing Korean-Language LLM Political Bias via Census-Grounded Agent Simulation
di: Kang, Sungwoo
Pubblicazione: (2026)
di: Kang, Sungwoo
Pubblicazione: (2026)
How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
di: Li, Yuxuan, et al.
Pubblicazione: (2026)
di: Li, Yuxuan, et al.
Pubblicazione: (2026)
MPO: Boosting LLM Agents with Meta Plan Optimization
di: Xiong, Weimin, et al.
Pubblicazione: (2025)
di: Xiong, Weimin, et al.
Pubblicazione: (2025)
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses
di: Jiang, Pengcheng, et al.
Pubblicazione: (2026)
di: Jiang, Pengcheng, et al.
Pubblicazione: (2026)
Can Generative Agents Predict Emotion?
di: Regan, Ciaran, et al.
Pubblicazione: (2024)
di: Regan, Ciaran, et al.
Pubblicazione: (2024)
ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
di: Li, Zhigen, et al.
Pubblicazione: (2024)
di: Li, Zhigen, et al.
Pubblicazione: (2024)
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
di: Sinha, Aarush, et al.
Pubblicazione: (2026)
di: Sinha, Aarush, et al.
Pubblicazione: (2026)
Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases
di: Li, Dubai, et al.
Pubblicazione: (2026)
di: Li, Dubai, et al.
Pubblicazione: (2026)
REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?
di: Jiang, Chenxi, et al.
Pubblicazione: (2025)
di: Jiang, Chenxi, et al.
Pubblicazione: (2025)
Attention-MoA: Enhancing Mixture-of-Agents via Inter-Agent Semantic Attention and Deep Residual Synthesis
di: Wen, Jianyu, et al.
Pubblicazione: (2026)
di: Wen, Jianyu, et al.
Pubblicazione: (2026)
MIRIX: Multi-Agent Memory System for LLM-Based Agents
di: Wang, Yu, et al.
Pubblicazione: (2025)
di: Wang, Yu, et al.
Pubblicazione: (2025)
Revealing the Barriers of Language Agents in Planning
di: Xie, Jian, et al.
Pubblicazione: (2024)
di: Xie, Jian, et al.
Pubblicazione: (2024)
Documenti analoghi
-
How Much Can RAG Help the Reasoning of LLM?
di: Liu, Jingyu, et al.
Pubblicazione: (2024) -
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
di: Islam, Saad Obaid ul, et al.
Pubblicazione: (2025) -
Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective
di: Aghzal, Mohamed, et al.
Pubblicazione: (2026) -
Code as Agent Harness
di: Ning, Xuying, et al.
Pubblicazione: (2026) -
It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers
di: Cho, Yong-eun
Pubblicazione: (2026)