Decoupling Time and Risk: Risk-Sensitive Reinforcement Learning with General Discounting

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Moghimi, Mehrdad, Coache, Anthony, Ku, Hyejin
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911421173858304
author Moghimi, Mehrdad
Coache, Anthony
Ku, Hyejin
author_facet Moghimi, Mehrdad
Coache, Anthony
Ku, Hyejin
contents Distributional reinforcement learning (RL) is a powerful framework increasingly adopted in safety-critical domains for its ability to optimize risk-sensitive objectives. However, the role of the discount factor is often overlooked, as it is typically treated as a fixed parameter of the Markov decision process or tunable hyperparameter, with little consideration of its effect on the learned policy. In the literature, it is well-known that the discounting function plays a major role in characterizing time preferences of an agent, which an exponential discount factor cannot fully capture. Building on this insight, we propose a novel framework that supports flexible discounting of future rewards and optimization of risk measures in distributional RL. We provide a technical analysis of the optimality of our algorithms, show that our multi-horizon extension fixes issues raised with existing methodologies, and validate the robustness of our methods through extensive experiments. Our results highlight that discounting is a cornerstone in decision-making problems for capturing more expressive temporal and risk preferences profiles, with potential implications for real-world safety-critical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04131
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Decoupling Time and Risk: Risk-Sensitive Reinforcement Learning with General Discounting
Moghimi, Mehrdad
Coache, Anthony
Ku, Hyejin
Machine Learning
Distributional reinforcement learning (RL) is a powerful framework increasingly adopted in safety-critical domains for its ability to optimize risk-sensitive objectives. However, the role of the discount factor is often overlooked, as it is typically treated as a fixed parameter of the Markov decision process or tunable hyperparameter, with little consideration of its effect on the learned policy. In the literature, it is well-known that the discounting function plays a major role in characterizing time preferences of an agent, which an exponential discount factor cannot fully capture. Building on this insight, we propose a novel framework that supports flexible discounting of future rewards and optimization of risk measures in distributional RL. We provide a technical analysis of the optimality of our algorithms, show that our multi-horizon extension fixes issues raised with existing methodologies, and validate the robustness of our methods through extensive experiments. Our results highlight that discounting is a cornerstone in decision-making problems for capturing more expressive temporal and risk preferences profiles, with potential implications for real-world safety-critical applications.
title Decoupling Time and Risk: Risk-Sensitive Reinforcement Learning with General Discounting
topic Machine Learning
url https://arxiv.org/abs/2602.04131