Saved in:
Bibliographic Details
Main Authors: Chen, Xiaoliang, Yu, Xin, Chang, Le, Jing, Teng, He, Jiashuai, Wang, Ze, Luo, Yangjun, Chen, Xingyu, Liang, Jiayue, Wang, Yuchen, Xie, Jiaying
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.18653
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909753339281408
author Chen, Xiaoliang
Yu, Xin
Chang, Le
Jing, Teng
He, Jiashuai
Wang, Ze
Luo, Yangjun
Chen, Xingyu
Liang, Jiayue
Wang, Yuchen
Xie, Jiaying
author_facet Chen, Xiaoliang
Yu, Xin
Chang, Le
Jing, Teng
He, Jiashuai
Wang, Ze
Luo, Yangjun
Chen, Xingyu
Liang, Jiayue
Wang, Yuchen
Xie, Jiaying
contents Information asymmetry in financial markets, often amplified by strategically crafted corporate narratives, undermines the effectiveness of conventional textual analysis. We propose a novel multimodal framework for financial risk assessment that integrates textual sentiment with paralinguistic cues derived from executive vocal tract dynamics in earnings calls. Central to this framework is the Physics-Informed Acoustic Model (PIAM), which applies nonlinear acoustics to robustly extract emotional signatures from raw teleconference sound subject to distortions such as signal clipping. Both acoustic and textual emotional states are projected onto an interpretable three-dimensional Affective State Label (ASL) space-Tension, Stability, and Arousal. Using a dataset of 1,795 earnings calls (approximately 1,800 hours), we construct features capturing dynamic shifts in executive affect between scripted presentation and spontaneous Q&A exchanges. Our key finding reveals a pronounced divergence in predictive capacity: while multimodal features do not forecast directional stock returns, they explain up to 43.8% of the out-of-sample variance in 30-day realized volatility. Importantly, volatility predictions are strongly driven by emotional dynamics during executive transitions from scripted to spontaneous speech, particularly reduced textual stability and heightened acoustic instability from CFOs, and significant arousal variability from CEOs. An ablation study confirms that our multimodal approach substantially outperforms a financials-only baseline, underscoring the complementary contributions of acoustic and textual modalities. By decoding latent markers of uncertainty from verifiable biometric signals, our methodology provides investors and regulators a powerful tool for enhancing market interpretability and identifying hidden corporate uncertainty.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18653
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability
Chen, Xiaoliang
Yu, Xin
Chang, Le
Jing, Teng
He, Jiashuai
Wang, Ze
Luo, Yangjun
Chen, Xingyu
Liang, Jiayue
Wang, Yuchen
Xie, Jiaying
Machine Learning
Artificial Intelligence
Sound
Audio and Speech Processing
62P05, 68T0
I.2.7; J.4
Information asymmetry in financial markets, often amplified by strategically crafted corporate narratives, undermines the effectiveness of conventional textual analysis. We propose a novel multimodal framework for financial risk assessment that integrates textual sentiment with paralinguistic cues derived from executive vocal tract dynamics in earnings calls. Central to this framework is the Physics-Informed Acoustic Model (PIAM), which applies nonlinear acoustics to robustly extract emotional signatures from raw teleconference sound subject to distortions such as signal clipping. Both acoustic and textual emotional states are projected onto an interpretable three-dimensional Affective State Label (ASL) space-Tension, Stability, and Arousal. Using a dataset of 1,795 earnings calls (approximately 1,800 hours), we construct features capturing dynamic shifts in executive affect between scripted presentation and spontaneous Q&A exchanges. Our key finding reveals a pronounced divergence in predictive capacity: while multimodal features do not forecast directional stock returns, they explain up to 43.8% of the out-of-sample variance in 30-day realized volatility. Importantly, volatility predictions are strongly driven by emotional dynamics during executive transitions from scripted to spontaneous speech, particularly reduced textual stability and heightened acoustic instability from CFOs, and significant arousal variability from CEOs. An ablation study confirms that our multimodal approach substantially outperforms a financials-only baseline, underscoring the complementary contributions of acoustic and textual modalities. By decoding latent markers of uncertainty from verifiable biometric signals, our methodology provides investors and regulators a powerful tool for enhancing market interpretability and identifying hidden corporate uncertainty.
title The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability
topic Machine Learning
Artificial Intelligence
Sound
Audio and Speech Processing
62P05, 68T0
I.2.7; J.4
url https://arxiv.org/abs/2508.18653