TraderBench: How Robust Are AI Agents in Adversarial Capital Markets?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Xiaochuang, Xu, Hui, Xu, Silvia, Zou, Cui, Xiong, Jing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918361287360512
author Yuan, Xiaochuang
Xu, Hui
Xu, Silvia
Zou, Cui
Xiong, Jing
author_facet Yuan, Xiaochuang
Xu, Hui
Xu, Silvia
Zou, Cui
Xiong, Jing
contents Evaluating AI agents in finance faces two key challenges: static benchmarks require costly expert annotation yet miss the dynamic decision-making central to real-world trading, while LLM-based judges introduce uncontrolled variance on domain-specific tasks. We introduce TraderBench, a benchmark that addresses both issues. It combines expert-verified static tasks (knowledge retrieval, analytical reasoning) with adversarial trading simulations scored purely on realized performance-Sharpe ratio, returns, and drawdown-eliminating judge variance entirely. The framework features two novel tracks: crypto trading with four progressive market-manipulation transforms, and options derivatives scoring across P&L accuracy, Greeks, and risk management. Trading scenarios can be refreshed with new market data to prevent benchmark contamination. Evaluating 13 models (8B open-source to frontier) on ~50 tasks, we find: (1) 8 of 13 models score ~33 on crypto with <1-point variation across adversarial conditions, exposing fixed non-adaptive strategies; (2) extended thinking helps retrieval (+26 points) but has zero impact on trading (+0.3 crypto, -0.1 options). These findings reveal that current agents lack genuine market adaptation, underscoring the need for performance-grounded evaluation in finance.
format Preprint
id arxiv_https___arxiv_org_abs_2603_00285
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TraderBench: How Robust Are AI Agents in Adversarial Capital Markets?
Yuan, Xiaochuang
Xu, Hui
Xu, Silvia
Zou, Cui
Xiong, Jing
Artificial Intelligence
I.2.11; J.4
Evaluating AI agents in finance faces two key challenges: static benchmarks require costly expert annotation yet miss the dynamic decision-making central to real-world trading, while LLM-based judges introduce uncontrolled variance on domain-specific tasks. We introduce TraderBench, a benchmark that addresses both issues. It combines expert-verified static tasks (knowledge retrieval, analytical reasoning) with adversarial trading simulations scored purely on realized performance-Sharpe ratio, returns, and drawdown-eliminating judge variance entirely. The framework features two novel tracks: crypto trading with four progressive market-manipulation transforms, and options derivatives scoring across P&L accuracy, Greeks, and risk management. Trading scenarios can be refreshed with new market data to prevent benchmark contamination. Evaluating 13 models (8B open-source to frontier) on ~50 tasks, we find: (1) 8 of 13 models score ~33 on crypto with <1-point variation across adversarial conditions, exposing fixed non-adaptive strategies; (2) extended thinking helps retrieval (+26 points) but has zero impact on trading (+0.3 crypto, -0.1 options). These findings reveal that current agents lack genuine market adaptation, underscoring the need for performance-grounded evaluation in finance.
title TraderBench: How Robust Are AI Agents in Adversarial Capital Markets?
topic Artificial Intelligence
I.2.11; J.4
url https://arxiv.org/abs/2603.00285