From Hypotheses to Factors: Constrained LLM Agents in Cryptocurrency Markets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yikuan, Fan, Zheqi, Hu, Kaiqi, Ye, Yifan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909000849686528
author Huang, Yikuan
Fan, Zheqi
Hu, Kaiqi
Ye, Yifan
author_facet Huang, Yikuan
Fan, Zheqi
Hu, Kaiqi
Ye, Yifan
contents LLM agents are promising tools for empirical discovery, but their flexibility can also turn discovery into uncontrolled search. We study how to use agents under a reproducible protocol through cryptocurrency factor discovery. Our framework casts the task as sequential hypothesis search: an agent reads an append-only experiment trace, proposes falsifiable factor hypotheses, and maps them to executable recipes, while a deterministic engine enforces fixed data splits, selection gates, transaction costs, and portfolio tests. Candidate actions are restricted to a point-in-time factor DSL, making both successful and failed hypotheses auditable. A ridge-combined portfolio trained only on 2020--2022 data achieves a 44.55% annualized return and Sharpe ratio of 1.55 in the 2024--2026 pure out-of-sample period after a 5 basis point one-way trading cost.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26747
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Hypotheses to Factors: Constrained LLM Agents in Cryptocurrency Markets
Huang, Yikuan
Fan, Zheqi
Hu, Kaiqi
Ye, Yifan
Portfolio Management
General Finance
Trading and Market Microstructure
LLM agents are promising tools for empirical discovery, but their flexibility can also turn discovery into uncontrolled search. We study how to use agents under a reproducible protocol through cryptocurrency factor discovery. Our framework casts the task as sequential hypothesis search: an agent reads an append-only experiment trace, proposes falsifiable factor hypotheses, and maps them to executable recipes, while a deterministic engine enforces fixed data splits, selection gates, transaction costs, and portfolio tests. Candidate actions are restricted to a point-in-time factor DSL, making both successful and failed hypotheses auditable. A ridge-combined portfolio trained only on 2020--2022 data achieves a 44.55% annualized return and Sharpe ratio of 1.55 in the 2024--2026 pure out-of-sample period after a 5 basis point one-way trading cost.
title From Hypotheses to Factors: Constrained LLM Agents in Cryptocurrency Markets
topic Portfolio Management
General Finance
Trading and Market Microstructure
url https://arxiv.org/abs/2604.26747