TIDE: Trajectory-based Diagnostic Evaluation of Test-Time Improvement in LLM Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yan, Hang, Che, Xinyu, Xu, Fangzhi, Sun, Qiushi, Ding, Zichen, Cheng, Kanzhi, Zhang, Jian, Qin, Tao, Liu, Jun, Lin, Qika
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917243585036288
author Yan, Hang
Che, Xinyu
Xu, Fangzhi
Sun, Qiushi
Ding, Zichen
Cheng, Kanzhi
Zhang, Jian
Qin, Tao
Liu, Jun
Lin, Qika
author_facet Yan, Hang
Che, Xinyu
Xu, Fangzhi
Sun, Qiushi
Ding, Zichen
Cheng, Kanzhi
Zhang, Jian
Qin, Tao
Liu, Jun
Lin, Qika
contents Recent advances in autonomous LLM agents demonstrate their ability to improve performance through iterative interaction with the environment. We define this paradigm as Test-Time Improvement (TTI). However, the mechanisms under how and why TTI succeed or fail remain poorly understood, and existing evaluation metrics fail to capture their task optimization efficiency, behavior adaptation after erroneous actions, and the specific utility of working memory for task completion. To address these gaps, we propose Test-time Improvement Diagnostic Evaluation (TIDE), an agent-agnostic and environment-agnostic framework that decomposes TTI into three comprehensive and interconnected dimensions. The framework measures (1) the overall temporal dynamics of task completion and (2) identifies whether performance is primarily constrained by recursive looping behaviors or (3) by burdensome accumulated memory. Through extensive experiments across diverse agents and environments, TIDE highlights that improving agent performance requires more than scaling internal reasoning, calling for explicitly optimizing the interaction dynamics between the agent and the environment.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02196
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TIDE: Trajectory-based Diagnostic Evaluation of Test-Time Improvement in LLM Agents
Yan, Hang
Che, Xinyu
Xu, Fangzhi
Sun, Qiushi
Ding, Zichen
Cheng, Kanzhi
Zhang, Jian
Qin, Tao
Liu, Jun
Lin, Qika
Artificial Intelligence
Recent advances in autonomous LLM agents demonstrate their ability to improve performance through iterative interaction with the environment. We define this paradigm as Test-Time Improvement (TTI). However, the mechanisms under how and why TTI succeed or fail remain poorly understood, and existing evaluation metrics fail to capture their task optimization efficiency, behavior adaptation after erroneous actions, and the specific utility of working memory for task completion. To address these gaps, we propose Test-time Improvement Diagnostic Evaluation (TIDE), an agent-agnostic and environment-agnostic framework that decomposes TTI into three comprehensive and interconnected dimensions. The framework measures (1) the overall temporal dynamics of task completion and (2) identifies whether performance is primarily constrained by recursive looping behaviors or (3) by burdensome accumulated memory. Through extensive experiments across diverse agents and environments, TIDE highlights that improving agent performance requires more than scaling internal reasoning, calling for explicitly optimizing the interaction dynamics between the agent and the environment.
title TIDE: Trajectory-based Diagnostic Evaluation of Test-Time Improvement in LLM Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2602.02196