A New DAPO Algorithm for Stock Trading

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zha, Ruijian, Liu, Bojun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916756164968448
author Zha, Ruijian
Liu, Bojun
author_facet Zha, Ruijian
Liu, Bojun
contents Recent advances in reinforcement learning, such as Dynamic Sampling Policy Optimization (DAPO), show strong performance when paired with large language models (LLMs). Motivated by this success, we ask whether similar gains can be realized in financial trading. We design a trading agent that combines an improved Group Relative Policy Optimization (GRPO) algorithm, augmented with ideas from DAPO, with LLM-based risk and sentiment signals extracted from financial news. On the NASDAQ-100 index (FNSPID dataset), our agent attains a cumulative return of 230.49 percent and an information ratio of 0.37, outperforming the CPPO-DeepSeek baseline. It also cuts training time from about 8 hours to 2.5 hours over 100 epochs while markedly reducing RAM usage. The proposed RL-LLM framework offers a scalable path toward data-efficient trading agents. Code: https://github.com/Ruijian-Zha/FinRL-DAPO-SR/
format Preprint
id arxiv_https___arxiv_org_abs_2505_06408
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A New DAPO Algorithm for Stock Trading
Zha, Ruijian
Liu, Bojun
Computational Engineering, Finance, and Science
68T05, 91G60
I.2.6; J.4; G.3
Recent advances in reinforcement learning, such as Dynamic Sampling Policy Optimization (DAPO), show strong performance when paired with large language models (LLMs). Motivated by this success, we ask whether similar gains can be realized in financial trading. We design a trading agent that combines an improved Group Relative Policy Optimization (GRPO) algorithm, augmented with ideas from DAPO, with LLM-based risk and sentiment signals extracted from financial news. On the NASDAQ-100 index (FNSPID dataset), our agent attains a cumulative return of 230.49 percent and an information ratio of 0.37, outperforming the CPPO-DeepSeek baseline. It also cuts training time from about 8 hours to 2.5 hours over 100 epochs while markedly reducing RAM usage. The proposed RL-LLM framework offers a scalable path toward data-efficient trading agents. Code: https://github.com/Ruijian-Zha/FinRL-DAPO-SR/
title A New DAPO Algorithm for Stock Trading
topic Computational Engineering, Finance, and Science
68T05, 91G60
I.2.6; J.4; G.3
url https://arxiv.org/abs/2505.06408