From Text to Alpha: Can LLMs Track Evolving Signals in Corporate Disclosures?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Choi, Chanyeol, Kim, Yoon, Yu, Yu, Cha, Young, Golkhou, V. Zach, Halperin, Igor, Papaioannou, Georgios, Kim, Minkyu, Wang, Zhangyang, Kwon, Jihoon, Kim, Minjae, Lopez-Lira, Alejandro, Lee, Yongjae
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914394423689216
author Choi, Chanyeol
Kim, Yoon
Yu, Yu
Cha, Young
Golkhou, V. Zach
Halperin, Igor
Papaioannou, Georgios
Kim, Minkyu
Wang, Zhangyang
Kwon, Jihoon
Kim, Minjae
Lopez-Lira, Alejandro
Lee, Yongjae
author_facet Choi, Chanyeol
Kim, Yoon
Yu, Yu
Cha, Young
Golkhou, V. Zach
Halperin, Igor
Papaioannou, Georgios
Kim, Minkyu
Wang, Zhangyang
Kwon, Jihoon
Kim, Minjae
Lopez-Lira, Alejandro
Lee, Yongjae
contents Natural language processing (NLP) has been widely used in quantitative finance, but traditional methods often struggle to capture rich narratives in corporate disclosures, leaving potentially informative signals under-explored. Large language models (LLMs) offer a promising alternative due to their ability to extract nuanced semantics. In this paper, we ask whether semantic signals extracted by LLMs from corporate disclosures predict alpha, defined as abnormal returns beyond broad market movements and common risk factors. We introduce a simple framework, LLM as extractor, embedding as ruler, which extracts context-aware, metric-focused textual spans and quantifies semantic changes across consecutive disclosure periods using embedding-based similarity. This allows us to measure the degree of metric shifting -- how much firms move away from previously emphasized metrics, referred as moving targets. In experiments with portfolio and cross-sectional regression tests against a recent NER-based baseline, our method achieves more than twice the risk-adjusted alpha and shows significantly stronger predictive power. Qualitative analysis suggests that these gains stem from preserving contextual qualifiers and filtering out non-metric terms that keyword-based approaches often miss.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03195
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Text to Alpha: Can LLMs Track Evolving Signals in Corporate Disclosures?
Choi, Chanyeol
Kim, Yoon
Yu, Yu
Cha, Young
Golkhou, V. Zach
Halperin, Igor
Papaioannou, Georgios
Kim, Minkyu
Wang, Zhangyang
Kwon, Jihoon
Kim, Minjae
Lopez-Lira, Alejandro
Lee, Yongjae
Computational Engineering, Finance, and Science
Natural language processing (NLP) has been widely used in quantitative finance, but traditional methods often struggle to capture rich narratives in corporate disclosures, leaving potentially informative signals under-explored. Large language models (LLMs) offer a promising alternative due to their ability to extract nuanced semantics. In this paper, we ask whether semantic signals extracted by LLMs from corporate disclosures predict alpha, defined as abnormal returns beyond broad market movements and common risk factors. We introduce a simple framework, LLM as extractor, embedding as ruler, which extracts context-aware, metric-focused textual spans and quantifies semantic changes across consecutive disclosure periods using embedding-based similarity. This allows us to measure the degree of metric shifting -- how much firms move away from previously emphasized metrics, referred as moving targets. In experiments with portfolio and cross-sectional regression tests against a recent NER-based baseline, our method achieves more than twice the risk-adjusted alpha and shows significantly stronger predictive power. Qualitative analysis suggests that these gains stem from preserving contextual qualifiers and filtering out non-metric terms that keyword-based approaches often miss.
title From Text to Alpha: Can LLMs Track Evolving Signals in Corporate Disclosures?
topic Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2510.03195