Measures of Variability for Risk-averse Policy Gradient
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Yudong, Pan, Yangchen, Tan, Jiaqi, Poupart, Pascal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Simple Mixture Policy Parameterization for Improving Sample Efficiency of CVaR Optimization
by: Luo, Yudong, et al.
Published: (2024)
by: Luo, Yudong, et al.
Published: (2024)
Gradient Residual Connections
by: Pan, Yangchen, et al.
Published: (2026)
by: Pan, Yangchen, et al.
Published: (2026)
Why Online Reinforcement Learning is Causal
by: Schulte, Oliver, et al.
Published: (2024)
by: Schulte, Oliver, et al.
Published: (2024)
The Reciprocity Gradient
by: Lin, Yue, et al.
Published: (2026)
by: Lin, Yue, et al.
Published: (2026)
TDHook: A Lightweight Framework for Interpretability
by: Poupart, Yoann
Published: (2025)
by: Poupart, Yoann
Published: (2025)
Reinforcement Learning in Dynamic Treatment Regimes Needs Critical Reexamination
by: Luo, Zhiyao, et al.
Published: (2024)
by: Luo, Zhiyao, et al.
Published: (2024)
Optimizing Risk-averse Human-AI Hybrid Teams
by: Fuchs, Andrew, et al.
Published: (2024)
by: Fuchs, Andrew, et al.
Published: (2024)
Risk-averse Total-reward MDPs with ERM and EVaR
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens
by: Jeong, Jihwan, et al.
Published: (2025)
by: Jeong, Jihwan, et al.
Published: (2025)
An MRP Formulation for Supervised Learning: Generalized Temporal Difference Learning Models
by: Pan, Yangchen, et al.
Published: (2024)
by: Pan, Yangchen, et al.
Published: (2024)
A Comprehensive Survey on Inverse Constrained Reinforcement Learning: Definitions, Progress and Challenges
by: Liu, Guiliang, et al.
Published: (2024)
by: Liu, Guiliang, et al.
Published: (2024)
DTR-Bench: An in silico Environment and Benchmark Platform for Reinforcement Learning Based Dynamic Treatment Regime
by: Luo, Zhiyao, et al.
Published: (2024)
by: Luo, Zhiyao, et al.
Published: (2024)
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024)
by: Pan, Deng, et al.
Published: (2024)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
by: Meng, Wenjia, et al.
Published: (2024)
by: Meng, Wenjia, et al.
Published: (2024)
CF-CAM: Cluster Filter Class Activation Mapping for Reliable Gradient-Based Interpretability
by: He, Hongjie, et al.
Published: (2025)
by: He, Hongjie, et al.
Published: (2025)
A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models
by: Miao, Yanting, et al.
Published: (2025)
by: Miao, Yanting, et al.
Published: (2025)
Complexity-Aware Deep Symbolic Regression with Robust Risk-Seeking Policy Gradients
by: Bastiani, Zachary, et al.
Published: (2024)
by: Bastiani, Zachary, et al.
Published: (2024)
LoRA-One: One-Step Full Gradient Could Suffice for Fine-Tuning Large Language Models, Provably and Efficiently
by: Zhang, Yuanhe, et al.
Published: (2025)
by: Zhang, Yuanhe, et al.
Published: (2025)
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025)
by: Ståhlberg, Simon, et al.
Published: (2025)
Policy Gradient with Kernel Quadrature
by: Hayakawa, Satoshi, et al.
Published: (2023)
by: Hayakawa, Satoshi, et al.
Published: (2023)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
Learning to Negotiate via Voluntary Commitment
by: Zhu, Shuhui, et al.
Published: (2025)
by: Zhu, Shuhui, et al.
Published: (2025)
Towards Efficient Risk-Sensitive Policy Gradient: An Iteration Complexity Analysis
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Label Alignment Regularization for Distribution Shift
by: Imani, Ehsan, et al.
Published: (2022)
by: Imani, Ehsan, et al.
Published: (2022)
Gradient Extrapolation-Based Policy Optimization
by: Swapnil, Ismam Nur, et al.
Published: (2026)
by: Swapnil, Ismam Nur, et al.
Published: (2026)
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
Randomness and Interpolation Improve Gradient Descent
by: Li, Jiawen, et al.
Published: (2025)
by: Li, Jiawen, et al.
Published: (2025)
Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning
by: Miao, Yanting, et al.
Published: (2024)
by: Miao, Yanting, et al.
Published: (2024)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024)
by: Xiao, Minheng, et al.
Published: (2024)
Gradient Regularized Natural Gradients
by: Dash, Satya Prakash, et al.
Published: (2026)
by: Dash, Satya Prakash, et al.
Published: (2026)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
Sequential Policy Gradient for Adaptive Hyperparameter Optimization
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
Soft Deterministic Policy Gradient with Gaussian Smoothing
by: Na, Hyunjun, et al.
Published: (2026)
by: Na, Hyunjun, et al.
Published: (2026)
Towards Provable Log Density Policy Gradient
by: Katdare, Pulkit, et al.
Published: (2024)
by: Katdare, Pulkit, et al.
Published: (2024)
Policy Gradient for Robust Markov Decision Processes
by: Wang, Qiuhao, et al.
Published: (2024)
by: Wang, Qiuhao, et al.
Published: (2024)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
by: Onoda, Ku, et al.
Published: (2026)
by: Onoda, Ku, et al.
Published: (2026)
Learning Soft Driving Constraints from Vectorized Scene Embeddings while Imitating Expert Trajectories
by: Mobarakeh, Niloufar Saeidi, et al.
Published: (2024)
by: Mobarakeh, Niloufar Saeidi, et al.
Published: (2024)
GRADE: Replacing Policy Gradients with Backpropagation for LLM Alignment
by: Nel, Lukas Abrie
Published: (2025)
by: Nel, Lukas Abrie
Published: (2025)
Vertical Symbolic Regression via Deep Policy Gradient
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
Policy Gradient Methods for Non-Markovian Reinforcement Learning
by: Kar, Avik, et al.
Published: (2026)
by: Kar, Avik, et al.
Published: (2026)
Similar Items
-
A Simple Mixture Policy Parameterization for Improving Sample Efficiency of CVaR Optimization
by: Luo, Yudong, et al.
Published: (2024) -
Gradient Residual Connections
by: Pan, Yangchen, et al.
Published: (2026) -
Why Online Reinforcement Learning is Causal
by: Schulte, Oliver, et al.
Published: (2024) -
The Reciprocity Gradient
by: Lin, Yue, et al.
Published: (2026) -
TDHook: A Lightweight Framework for Interpretability
by: Poupart, Yoann
Published: (2025)