Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Matrenok, Simon, Moalla, Skander, Gulcehre, Caglar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
von: Moalla, Skander, et al.
Veröffentlicht: (2024)
von: Moalla, Skander, et al.
Veröffentlicht: (2024)
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
von: Wei, Xiuying, et al.
Veröffentlicht: (2024)
Partition Generative Modeling: Masked Modeling Without Masks
von: Deschenaux, Justin, et al.
Veröffentlicht: (2025)
von: Deschenaux, Justin, et al.
Veröffentlicht: (2025)
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
von: Terekhov, Mikhail, et al.
Veröffentlicht: (2024)
von: Terekhov, Mikhail, et al.
Veröffentlicht: (2024)
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
BlockGen: Flexible Blockwise Sequence Modeling with Hybrid Samplers
von: Deschenaux, Justin, et al.
Veröffentlicht: (2026)
von: Deschenaux, Justin, et al.
Veröffentlicht: (2026)
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
von: Wei, Xiuying, et al.
Veröffentlicht: (2026)
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
von: Deschenaux, Justin, et al.
Veröffentlicht: (2024)
von: Deschenaux, Justin, et al.
Veröffentlicht: (2024)
Python Machine Learning Research Template
von: Moalla, Skander
Veröffentlicht: (2025)
von: Moalla, Skander
Veröffentlicht: (2025)
Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards
von: Karzanov, Daniil, et al.
Veröffentlicht: (2025)
von: Karzanov, Daniil, et al.
Veröffentlicht: (2025)
The Role of Deep Learning Regularizations on Actors in Offline RL
von: Tarasov, Denis, et al.
Veröffentlicht: (2024)
von: Tarasov, Denis, et al.
Veröffentlicht: (2024)
The Diffusion Duality, Chapter II: $Ψ$-Samplers
von: Deschenaux, Justin, et al.
Veröffentlicht: (2026)
von: Deschenaux, Justin, et al.
Veröffentlicht: (2026)
Value-Free Policy Optimization via Reward Partitioning
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
Quantile Regression for Distributional Reward Models in RLHF
von: Dorka, Nicolai
Veröffentlicht: (2024)
von: Dorka, Nicolai
Veröffentlicht: (2024)
Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
von: Jo, Mingyu, et al.
Veröffentlicht: (2025)
von: Jo, Mingyu, et al.
Veröffentlicht: (2025)
HiPPO-Prophecy: State-Space Models can Provably Learn Dynamical Systems in Context
von: Joseph, Federico Arangath, et al.
Veröffentlicht: (2024)
von: Joseph, Federico Arangath, et al.
Veröffentlicht: (2024)
Distributional Off-Policy Evaluation with Deep Quantile Process Regression
von: Kuang, Qi, et al.
Veröffentlicht: (2026)
von: Kuang, Qi, et al.
Veröffentlicht: (2026)
Control Tax: The Price of Keeping AI in Check
von: Terekhov, Mikhail, et al.
Veröffentlicht: (2025)
von: Terekhov, Mikhail, et al.
Veröffentlicht: (2025)
Simple Hierarchical Planning with Diffusion
von: Chen, Chang, et al.
Veröffentlicht: (2024)
von: Chen, Chang, et al.
Veröffentlicht: (2024)
Boosting CVaR Policy Optimization with Quantile Gradients
von: Luo, Yudong, et al.
Veröffentlicht: (2026)
von: Luo, Yudong, et al.
Veröffentlicht: (2026)
Vector Quantile Regression on Manifolds
von: Pegoraro, Marco, et al.
Veröffentlicht: (2023)
von: Pegoraro, Marco, et al.
Veröffentlicht: (2023)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
von: Orvieto, Antonio, et al.
Veröffentlicht: (2023)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2023)
Quantile Q-Learning: Revisiting Offline Extreme Q-Learning with Quantile Regression
von: Gao, Xinming, et al.
Veröffentlicht: (2025)
von: Gao, Xinming, et al.
Veröffentlicht: (2025)
An Efficient Multi Quantile Regression Network with Ad Hoc Prevention of Quantile Crossing
von: Decke, Jens, et al.
Veröffentlicht: (2024)
von: Decke, Jens, et al.
Veröffentlicht: (2024)
Multi-Fidelity Quantile Regression
von: Liu, Yixiang, et al.
Veröffentlicht: (2026)
von: Liu, Yixiang, et al.
Veröffentlicht: (2026)
PlanDQ: Hierarchical Plan Orchestration via D-Conductor and Q-Performer
von: Chen, Chang, et al.
Veröffentlicht: (2024)
von: Chen, Chang, et al.
Veröffentlicht: (2024)
An Exact Pointwise Characterization for Total Variation Denoising in Quantile Regression
von: Ghoshal, Deep, et al.
Veröffentlicht: (2026)
von: Ghoshal, Deep, et al.
Veröffentlicht: (2026)
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
von: Sahu, Sharan, et al.
Veröffentlicht: (2025)
von: Sahu, Sharan, et al.
Veröffentlicht: (2025)
Horseshoe Prior Bayesian Quantile Regression
von: Kohns, David, et al.
Veröffentlicht: (2020)
von: Kohns, David, et al.
Veröffentlicht: (2020)
Interpretable Quantile Regression by Optimal Decision Trees
von: Lemaire, Valentin, et al.
Veröffentlicht: (2026)
von: Lemaire, Valentin, et al.
Veröffentlicht: (2026)
ReModels: Quantile Regression Averaging models
von: Zakrzewski, Grzegorz, et al.
Veröffentlicht: (2024)
von: Zakrzewski, Grzegorz, et al.
Veröffentlicht: (2024)
Symbolic Quantile Regression for the Interpretable Prediction of Conditional Quantiles
von: Hoekstra, Cas Oude, et al.
Veröffentlicht: (2025)
von: Hoekstra, Cas Oude, et al.
Veröffentlicht: (2025)
On Learning the Tail Quantiles of Driving Behavior Distributions via Quantile Regression and Flows
von: Tee, Jia Yu, et al.
Veröffentlicht: (2023)
von: Tee, Jia Yu, et al.
Veröffentlicht: (2023)
On the Pointwise Behavior of Recursive Partitioning and Its Implications for Heterogeneous Causal Effect Estimation
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2022)
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2022)
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
von: Song, Haonan, et al.
Veröffentlicht: (2026)
von: Song, Haonan, et al.
Veröffentlicht: (2026)
Integrating Uncertainty Awareness into Conformalized Quantile Regression
von: Rossellini, Raphael, et al.
Veröffentlicht: (2023)
von: Rossellini, Raphael, et al.
Veröffentlicht: (2023)
fastkqr: A Fast Algorithm for Kernel Quantile Regression
von: Tang, Qian, et al.
Veröffentlicht: (2024)
von: Tang, Qian, et al.
Veröffentlicht: (2024)
Spatial Conformal Inference through Localized Quantile Regression
von: Jiang, Hanyang, et al.
Veröffentlicht: (2024)
von: Jiang, Hanyang, et al.
Veröffentlicht: (2024)
Relaxed Quantile Regression: Prediction Intervals for Asymmetric Noise
von: Pouplin, Thomas, et al.
Veröffentlicht: (2024)
von: Pouplin, Thomas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
von: Moalla, Skander, et al.
Veröffentlicht: (2024) -
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
von: Wei, Xiuying, et al.
Veröffentlicht: (2024) -
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
von: Wei, Xiuying, et al.
Veröffentlicht: (2024) -
Partition Generative Modeling: Masked Modeling Without Masks
von: Deschenaux, Justin, et al.
Veröffentlicht: (2025) -
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
von: Terekhov, Mikhail, et al.
Veröffentlicht: (2024)