Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Jungsoo, Chae, Hyungjoo, Mendes, Ethan, DeYoung, Jay, Kishore, Varsha, Xu, Wei, Ritter, Alan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Safe and Scalable Web Agent Learning via Recreated Websites
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2026)
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2026)
Didactic to Constructive: Turning Expert Solutions into Learnable Reasoning
von: Mendes, Ethan, et al.
Veröffentlicht: (2026)
von: Mendes, Ethan, et al.
Veröffentlicht: (2026)
Anticipatory Evaluation of Language Models
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
Improving Attributed Long-form Question Answering with Intent Awareness
von: Zhao, Xinran, et al.
Veröffentlicht: (2026)
von: Zhao, Xinran, et al.
Veröffentlicht: (2026)
Language Models can Self-Improve at State-Value Estimation for Better Search
von: Mendes, Ethan, et al.
Veröffentlicht: (2025)
von: Mendes, Ethan, et al.
Veröffentlicht: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
Rural Education: Issues and Practice. Source Books on Education, Vol. 25. Garland Reference Library of Social Science, Vol. 473.
von: DeYoung, Alan J., Ed.
Veröffentlicht: (1991)
von: DeYoung, Alan J., Ed.
Veröffentlicht: (1991)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
von: Kim, Sunghwan, et al.
Veröffentlicht: (2025)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2025)
Resource Sharing of Micro-Software, or, What Ever Happened to All That CP/M Compatibility?
von: DeYoung, Barbara
Veröffentlicht: (1984)
von: DeYoung, Barbara
Veröffentlicht: (1984)
Distribution-Aware Reward Estimation for Test-Time Reinforcement Learning
von: Du, Bodong, et al.
Veröffentlicht: (2026)
von: Du, Bodong, et al.
Veröffentlicht: (2026)
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
von: Sahu, Sharan, et al.
Veröffentlicht: (2025)
von: Sahu, Sharan, et al.
Veröffentlicht: (2025)
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
von: Du, Yao, et al.
Veröffentlicht: (2026)
von: Du, Yao, et al.
Veröffentlicht: (2026)
Granular Privacy Control for Geolocation with Vision Language Models
von: Mendes, Ethan, et al.
Veröffentlicht: (2024)
von: Mendes, Ethan, et al.
Veröffentlicht: (2024)
Investigating and Alleviating Harm Amplification in LLM Interactions
von: Guo, Ruohao, et al.
Veröffentlicht: (2026)
von: Guo, Ruohao, et al.
Veröffentlicht: (2026)
Quantile Regression for Distributional Reward Models in RLHF
von: Dorka, Nicolai
Veröffentlicht: (2024)
von: Dorka, Nicolai
Veröffentlicht: (2024)
The Distributional Reward Critic Framework for Reinforcement Learning Under Perturbed Rewards
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
Loss- and Reward-Weighting for Efficient Distributed Reinforcement Learning
von: Holen, Martin, et al.
Veröffentlicht: (2023)
von: Holen, Martin, et al.
Veröffentlicht: (2023)
Reward Augmentation in Reinforcement Learning for Testing Distributed Systems
von: Borgarelli, Andrea, et al.
Veröffentlicht: (2024)
von: Borgarelli, Andrea, et al.
Veröffentlicht: (2024)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
von: Du, Chengze, et al.
Veröffentlicht: (2025)
von: Du, Chengze, et al.
Veröffentlicht: (2025)
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
von: Guo, Ruohao, et al.
Veröffentlicht: (2025)
von: Guo, Ruohao, et al.
Veröffentlicht: (2025)
Airborne Bacterial Communities: Diversity, Survival Strategies and Functional Roles in the Atmosphere
von: Jungsoo Park, et al.
Veröffentlicht: (2026)
von: Jungsoo Park, et al.
Veröffentlicht: (2026)
Unexpected Shifts in Sleep Architecture Associated with Cognitive Impairment
von: Hyeonseul Park, et al.
Veröffentlicht: (2025)
von: Hyeonseul Park, et al.
Veröffentlicht: (2025)
Comparative Study of Key Reactions Affecting Nitrogen Oxide Formation in Methane Laminar Premixed Flame
von: Jonghyun Kim, et al.
Veröffentlicht: (2024)
von: Jonghyun Kim, et al.
Veröffentlicht: (2024)
Distributional Reinforcement Learning with Dual Expectile-Quantile Regression
von: Jullien, Sami, et al.
Veröffentlicht: (2023)
von: Jullien, Sami, et al.
Veröffentlicht: (2023)
Topology-Aware Reinforcement Learning over Graphs for Resilient Power Distribution Networks
von: Jacob, Roshni Anna, et al.
Veröffentlicht: (2026)
von: Jacob, Roshni Anna, et al.
Veröffentlicht: (2026)
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge
von: Zhang, Yasi, et al.
Veröffentlicht: (2026)
von: Zhang, Yasi, et al.
Veröffentlicht: (2026)
Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?
von: Yang, Ruixin, et al.
Veröffentlicht: (2026)
von: Yang, Ruixin, et al.
Veröffentlicht: (2026)
Trajectory Planning for Autonomous Vehicle Using Iterative Reward Prediction in Reinforcement Learning
von: Park, Hyunwoo
Veröffentlicht: (2024)
von: Park, Hyunwoo
Veröffentlicht: (2024)
Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement Learning
von: Rathbun, Ethan, et al.
Veröffentlicht: (2026)
von: Rathbun, Ethan, et al.
Veröffentlicht: (2026)
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning
von: Wang, Yujie, et al.
Veröffentlicht: (2026)
von: Wang, Yujie, et al.
Veröffentlicht: (2026)
REBEL: Reinforcement Learning via Regressing Relative Rewards
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks
von: Kim, Jeongmo, et al.
Veröffentlicht: (2025)
von: Kim, Jeongmo, et al.
Veröffentlicht: (2025)
Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning
von: Kaya, Ege C., et al.
Veröffentlicht: (2026)
von: Kaya, Ege C., et al.
Veröffentlicht: (2026)
Do Multi-Document Summarization Models Synthesize?
von: DeYoung, Jay, et al.
Veröffentlicht: (2023)
von: DeYoung, Jay, et al.
Veröffentlicht: (2023)
On the strong Arnold chord conjecture for convex contact forms
von: Kang, Jungsoo
Veröffentlicht: (2023)
von: Kang, Jungsoo
Veröffentlicht: (2023)
On the strong Arnold chord conjecture for convex contact forms
von: Jungsoo Kang
Veröffentlicht: (2026)
von: Jungsoo Kang
Veröffentlicht: (2026)
Ähnliche Einträge
-
Safe and Scalable Web Agent Learning via Recreated Websites
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2026) -
Didactic to Constructive: Turning Expert Solutions into Learnable Reasoning
von: Mendes, Ethan, et al.
Veröffentlicht: (2026) -
Anticipatory Evaluation of Language Models
von: Park, Jungsoo, et al.
Veröffentlicht: (2025) -
Evaluating Robustness of Reward Models for Mathematical Reasoning
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024) -
Improving Attributed Long-form Question Answering with Intent Awareness
von: Zhao, Xinran, et al.
Veröffentlicht: (2026)