A Theoretical Framework for Partially Observed Reward-States in RLHF
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kausik, Chinmaya, Mutti, Mirco, Pacchiano, Aldo, Tewari, Ambuj |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Offline Data in Linear Latent Contextual Bandits
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025)
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2026)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2026)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Second Order Bounds for Contextual Bandits with Function Approximation
von: Pacchiano, Aldo
Veröffentlicht: (2024)
von: Pacchiano, Aldo
Veröffentlicht: (2024)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
State-free Reinforcement Learning
von: Chen, Mingyu, et al.
Veröffentlicht: (2024)
von: Chen, Mingyu, et al.
Veröffentlicht: (2024)
Learning Rate-Free Reinforcement Learning: A Case for Model Selection with Non-Stationary Objectives
von: Afshar, Aida, et al.
Veröffentlicht: (2024)
von: Afshar, Aida, et al.
Veröffentlicht: (2024)
Compute Aligned Training: Optimizing for Test Time Inference
von: Ousherovitch, Adam, et al.
Veröffentlicht: (2026)
von: Ousherovitch, Adam, et al.
Veröffentlicht: (2026)
Improved Training Mechanism for Reinforcement Learning via Online Model Selection
von: Afshar, Aida, et al.
Veröffentlicht: (2025)
von: Afshar, Aida, et al.
Veröffentlicht: (2025)
How to Explore with Belief: State Entropy Maximization in POMDPs
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2024)
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2024)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
Information-Theoretic Reward Decomposition for Generalizable RLHF
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
von: Miao, Yuchun, et al.
Veröffentlicht: (2024)
von: Miao, Yuchun, et al.
Veröffentlicht: (2024)
Towards Principled Unsupervised Multi-Agent Reinforcement Learning
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2025)
von: Zamboni, Riccardo, et al.
Veröffentlicht: (2025)
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
von: Misra, Dipendra, et al.
Veröffentlicht: (2026)
von: Misra, Dipendra, et al.
Veröffentlicht: (2026)
Data-Driven Online Model Selection With Regret Guarantees
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2023)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2023)
In-Context Learning for Pure Exploration
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching
von: Fraschini, Andrea, et al.
Veröffentlicht: (2026)
von: Fraschini, Andrea, et al.
Veröffentlicht: (2026)
Multiple-policy Evaluation via Density Estimation
von: Chen, Yilei, et al.
Veröffentlicht: (2024)
von: Chen, Yilei, et al.
Veröffentlicht: (2024)
Experiment Planning with Function Approximation
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
von: Tenedini, Davide, et al.
Veröffentlicht: (2025)
von: Tenedini, Davide, et al.
Veröffentlicht: (2025)
Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF
von: Liu, Jing
Veröffentlicht: (2025)
von: Liu, Jing
Veröffentlicht: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Reward Compatibility: A Framework for Inverse RL
von: Lazzati, Filippo, et al.
Veröffentlicht: (2025)
von: Lazzati, Filippo, et al.
Veröffentlicht: (2025)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
von: Zhang, Chuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Chuheng, et al.
Veröffentlicht: (2024)
In-Context Learning for Pure Exploration in Continuous Spaces
von: Russo, Alessio, et al.
Veröffentlicht: (2026)
von: Russo, Alessio, et al.
Veröffentlicht: (2026)
Provable Interactive Learning with Hindsight Instruction Feedback
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
von: Baharav, Tavor Z., et al.
Veröffentlicht: (2025)
von: Baharav, Tavor Z., et al.
Veröffentlicht: (2025)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
A Characterization of List Language Identification in the Limit
von: Charikar, Moses, et al.
Veröffentlicht: (2025)
von: Charikar, Moses, et al.
Veröffentlicht: (2025)
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
von: Tang, Kenton, et al.
Veröffentlicht: (2026)
von: Tang, Kenton, et al.
Veröffentlicht: (2026)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
von: Zhao, Kangwen, et al.
Veröffentlicht: (2025)
von: Zhao, Kangwen, et al.
Veröffentlicht: (2025)
Quantile Regression for Distributional Reward Models in RLHF
von: Dorka, Nicolai
Veröffentlicht: (2024)
von: Dorka, Nicolai
Veröffentlicht: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Accelerating RLHF Training with Reward Variance Increase
von: Yang, Zonglin, et al.
Veröffentlicht: (2025)
von: Yang, Zonglin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Offline Data in Linear Latent Contextual Bandits
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024) -
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025) -
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2026) -
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
von: Russo, Alessio, et al.
Veröffentlicht: (2025) -
Second Order Bounds for Contextual Bandits with Function Approximation
von: Pacchiano, Aldo
Veröffentlicht: (2024)