Trade-R1: Bridging Verifiable Rewards to Stochastic Environments via Process-Level Reasoning Verification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Rui, Sun, Yifan, Xu, Sheng, Zhao, Li, Li, Jing, Jiang, Daxin, Hua, Cheng, Bai, Zuo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911359934922752
author Sun, Rui
Sun, Yifan
Xu, Sheng
Zhao, Li
Li, Jing
Jiang, Daxin
Hua, Cheng
Bai, Zuo
author_facet Sun, Rui
Sun, Yifan
Xu, Sheng
Zhao, Li
Li, Jing
Jiang, Daxin
Hua, Cheng
Bai, Zuo
contents Reinforcement Learning (RL) has enabled Large Language Models (LLMs) to achieve remarkable reasoning in domains like mathematics and coding, where verifiable rewards provide clear signals. However, extending this paradigm to financial decision is challenged by the market's stochastic nature: rewards are verifiable but inherently noisy, causing standard RL to degenerate into reward hacking. To address this, we propose Trade-R1, a model training framework that bridges verifiable rewards to stochastic environments via process-level reasoning verification. Our key innovation is a verification method that transforms the problem of evaluating reasoning over lengthy financial documents into a structured Retrieval-Augmented Generation (RAG) task. We construct a triangular consistency metric, assessing pairwise alignment between retrieved evidence, reasoning chains, and decisions to serve as a validity filter for noisy market returns. We explore two reward integration strategies: Fixed-effect Semantic Reward (FSR) for stable alignment signals, and Dynamic-effect Semantic Reward (DSR) for coupled magnitude optimization. Experiments on different country asset selection demonstrate that our paradigm reduces reward hacking, with DSR achieving superior cross-market generalization while maintaining the highest reasoning consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03948
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Trade-R1: Bridging Verifiable Rewards to Stochastic Environments via Process-Level Reasoning Verification
Sun, Rui
Sun, Yifan
Xu, Sheng
Zhao, Li
Li, Jing
Jiang, Daxin
Hua, Cheng
Bai, Zuo
Artificial Intelligence
Trading and Market Microstructure
Reinforcement Learning (RL) has enabled Large Language Models (LLMs) to achieve remarkable reasoning in domains like mathematics and coding, where verifiable rewards provide clear signals. However, extending this paradigm to financial decision is challenged by the market's stochastic nature: rewards are verifiable but inherently noisy, causing standard RL to degenerate into reward hacking. To address this, we propose Trade-R1, a model training framework that bridges verifiable rewards to stochastic environments via process-level reasoning verification. Our key innovation is a verification method that transforms the problem of evaluating reasoning over lengthy financial documents into a structured Retrieval-Augmented Generation (RAG) task. We construct a triangular consistency metric, assessing pairwise alignment between retrieved evidence, reasoning chains, and decisions to serve as a validity filter for noisy market returns. We explore two reward integration strategies: Fixed-effect Semantic Reward (FSR) for stable alignment signals, and Dynamic-effect Semantic Reward (DSR) for coupled magnitude optimization. Experiments on different country asset selection demonstrate that our paradigm reduces reward hacking, with DSR achieving superior cross-market generalization while maintaining the highest reasoning consistency.
title Trade-R1: Bridging Verifiable Rewards to Stochastic Environments via Process-Level Reasoning Verification
topic Artificial Intelligence
Trading and Market Microstructure
url https://arxiv.org/abs/2601.03948