AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Ruipeng, Chen, Yuxin, Wang, Yukai, Wu, Chang, Fang, Junfeng, Cai, Xiaodong, Gu, Qi, Su, Hui, Zhang, An, Wang, Xiang, Cai, Xunliang, Chua, Tat-Seng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912911113322496
author Wang, Ruipeng
Chen, Yuxin
Wang, Yukai
Wu, Chang
Fang, Junfeng
Cai, Xiaodong
Gu, Qi
Su, Hui
Zhang, An
Wang, Xiang
Cai, Xunliang
Chua, Tat-Seng
author_facet Wang, Ruipeng
Chen, Yuxin
Wang, Yukai
Wu, Chang
Fang, Junfeng
Cai, Xiaodong
Gu, Qi
Su, Hui
Zhang, An
Wang, Xiang
Cai, Xunliang
Chua, Tat-Seng
contents Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployments often that observed on benchmark settings, especially in complex and imperfect environments. This discrepancy largely arises because prevailing training and evaluation paradigms are typically built on idealized assumptions, overlooking the inherent stochasticity and noise present in real-world interactions. To bridge this gap, we introduce AgentNoiseBench, a framework for systematically evaluating the robustness of agentic models under noisy environments. We first conduct an in-depth analysis of biases and uncertainties in real-world scenarios and categorize environmental noise into two primary types: user-noise and tool-noise. Building on this analysis, we develop an automated pipeline that injects controllable noise into existing agent-centric benchmarks while preserving task solvability. Leveraging this pipeline, we perform extensive evaluations across a wide range of models with diverse architectures and parameter scales. Our results reveal consistent performance variations under different noise conditions, highlighting the sensitivity of current agentic models to realistic environmental perturbations.
format Preprint
id arxiv_https___arxiv_org_abs_2602_11348
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
Wang, Ruipeng
Chen, Yuxin
Wang, Yukai
Wu, Chang
Fang, Junfeng
Cai, Xiaodong
Gu, Qi
Su, Hui
Zhang, An
Wang, Xiang
Cai, Xunliang
Chua, Tat-Seng
Artificial Intelligence
Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployments often that observed on benchmark settings, especially in complex and imperfect environments. This discrepancy largely arises because prevailing training and evaluation paradigms are typically built on idealized assumptions, overlooking the inherent stochasticity and noise present in real-world interactions. To bridge this gap, we introduce AgentNoiseBench, a framework for systematically evaluating the robustness of agentic models under noisy environments. We first conduct an in-depth analysis of biases and uncertainties in real-world scenarios and categorize environmental noise into two primary types: user-noise and tool-noise. Building on this analysis, we develop an automated pipeline that injects controllable noise into existing agent-centric benchmarks while preserving task solvability. Leveraging this pipeline, we perform extensive evaluations across a wide range of models with diverse architectures and parameter scales. Our results reveal consistent performance variations under different noise conditions, highlighting the sensitivity of current agentic models to realistic environmental perturbations.
title AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
topic Artificial Intelligence
url https://arxiv.org/abs/2602.11348