It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers
Fuente:
arXiv
Saved in:
| Main Author: | Cho, Yong-eun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Many-Tier Instruction Hierarchy in LLM Agents
by: Zhang, Jingyu, et al.
Published: (2026)
by: Zhang, Jingyu, et al.
Published: (2026)
It's Not the Size: Harness Design Determines Operational Stability in Small Language Models
by: Cho, Yong-eun
Published: (2026)
by: Cho, Yong-eun
Published: (2026)
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
by: Tong, Zekai, et al.
Published: (2026)
by: Tong, Zekai, et al.
Published: (2026)
Code as Agent Harness
by: Ning, Xuying, et al.
Published: (2026)
by: Ning, Xuying, et al.
Published: (2026)
Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation
by: Yu, Ye, et al.
Published: (2026)
by: Yu, Ye, et al.
Published: (2026)
Automatic Transmission for LLM Tiers: Optimizing Cost and Accuracy in Large Language Models
by: Na, Injae, et al.
Published: (2025)
by: Na, Injae, et al.
Published: (2025)
Natural-Language Agent Harnesses
by: Pan, Linyue, et al.
Published: (2026)
by: Pan, Linyue, et al.
Published: (2026)
UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind
by: Qian, Cheng, et al.
Published: (2026)
by: Qian, Cheng, et al.
Published: (2026)
Personality Expression Across Contexts: Linguistic and Behavioral Variation in LLM Agents
by: Han, Bin, et al.
Published: (2026)
by: Han, Bin, et al.
Published: (2026)
How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent
by: Jung, Sungwoo, et al.
Published: (2026)
by: Jung, Sungwoo, et al.
Published: (2026)
ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild
by: Yang, Bufang, et al.
Published: (2025)
by: Yang, Bufang, et al.
Published: (2025)
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
by: Rahman, Mizanur, et al.
Published: (2025)
by: Rahman, Mizanur, et al.
Published: (2025)
WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
by: Lee, Woongkyu, et al.
Published: (2025)
by: Lee, Woongkyu, et al.
Published: (2025)
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
Non-Monotonic Attention-based Read/Write Policy Learning for Simultaneous Translation
by: Ahmed, Zeeshan, et al.
Published: (2025)
by: Ahmed, Zeeshan, et al.
Published: (2025)
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
Harnessing Consistency for Robust Test-Time LLM Ensemble
by: Zeng, Zhichen, et al.
Published: (2025)
by: Zeng, Zhichen, et al.
Published: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
by: Dzikanyanga, Gradwell, et al.
Published: (2026)
by: Dzikanyanga, Gradwell, et al.
Published: (2026)
Distilling LLM Agent into Small Models with Retrieval and Code Tools
by: Kang, Minki, et al.
Published: (2025)
by: Kang, Minki, et al.
Published: (2025)
Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory
by: Zhang, Haozhen, et al.
Published: (2026)
by: Zhang, Haozhen, et al.
Published: (2026)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025)
by: Pimpale, Govind, et al.
Published: (2025)
Stance Detection with Collaborative Role-Infused LLM-Based Agents
by: Lan, Xiaochong, et al.
Published: (2023)
by: Lan, Xiaochong, et al.
Published: (2023)
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
by: Han, Wenhan, et al.
Published: (2025)
by: Han, Wenhan, et al.
Published: (2025)
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses
by: Jiang, Pengcheng, et al.
Published: (2026)
by: Jiang, Pengcheng, et al.
Published: (2026)
LLM Rationalis? Measuring Bargaining Capabilities of AI Negotiators
by: Shah, Cheril, et al.
Published: (2025)
by: Shah, Cheril, et al.
Published: (2025)
Enforcing Monotonic Progress in Legal Cross-Examination: Preventing Long-Horizon Stagnation in LLM-Based Inquiry
by: Liao, Hsien-Jyh
Published: (2026)
by: Liao, Hsien-Jyh
Published: (2026)
VeRO: An Evaluation Harness for Agents to Optimize Agents
by: Ursekar, Varun, et al.
Published: (2026)
by: Ursekar, Varun, et al.
Published: (2026)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities
by: Chen, Jiaqi, et al.
Published: (2026)
by: Chen, Jiaqi, et al.
Published: (2026)
Self-Review Framework for Enhancing Instruction Following Capability of LLM
by: Park, Sihyun
Published: (2025)
by: Park, Sihyun
Published: (2025)
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities
by: Hong, Zhaochen, et al.
Published: (2025)
by: Hong, Zhaochen, et al.
Published: (2025)
ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning
by: Huang, Fan
Published: (2026)
by: Huang, Fan
Published: (2026)
AutoHarness: improving LLM agents by automatically synthesizing a code harness
by: Lou, Xinghua, et al.
Published: (2026)
by: Lou, Xinghua, et al.
Published: (2026)
LLM2: Let Large Language Models Harness System 2 Reasoning
by: Yang, Cheng, et al.
Published: (2024)
by: Yang, Cheng, et al.
Published: (2024)
A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device Agents
by: Shao, Chenyang, et al.
Published: (2025)
by: Shao, Chenyang, et al.
Published: (2025)
Harnessing Multi-Role Capabilities of Large Language Models for Open-Domain Question Answering
by: Sun, Hongda, et al.
Published: (2024)
by: Sun, Hongda, et al.
Published: (2024)
Detecting Safety Violations Across Many Agent Traces
by: Stein, Adam, et al.
Published: (2026)
by: Stein, Adam, et al.
Published: (2026)
Similar Items
-
Many-Tier Instruction Hierarchy in LLM Agents
by: Zhang, Jingyu, et al.
Published: (2026) -
It's Not the Size: Harness Design Determines Operational Stability in Small Language Models
by: Cho, Yong-eun
Published: (2026) -
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
by: Tong, Zekai, et al.
Published: (2026) -
Code as Agent Harness
by: Ning, Xuying, et al.
Published: (2026) -
Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation
by: Yu, Ye, et al.
Published: (2026)