A Theoretical Framework for Partially Observed Reward-States in RLHF
Fuente:
arXiv
Guardado en:
| Autores principales: | Kausik, Chinmaya, Mutti, Mirco, Pacchiano, Aldo, Tewari, Ambuj |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Leveraging Offline Data in Linear Latent Contextual Bandits
por: Kausik, Chinmaya, et al.
Publicado: (2024)
por: Kausik, Chinmaya, et al.
Publicado: (2024)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
por: Hong, Kihyuk, et al.
Publicado: (2025)
por: Hong, Kihyuk, et al.
Publicado: (2025)
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
por: Kausik, Chinmaya, et al.
Publicado: (2026)
por: Kausik, Chinmaya, et al.
Publicado: (2026)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
por: Russo, Alessio, et al.
Publicado: (2025)
por: Russo, Alessio, et al.
Publicado: (2025)
Second Order Bounds for Contextual Bandits with Function Approximation
por: Pacchiano, Aldo
Publicado: (2024)
por: Pacchiano, Aldo
Publicado: (2024)
Active Preference Optimization for Sample Efficient RLHF
por: Das, Nirjhar, et al.
Publicado: (2024)
por: Das, Nirjhar, et al.
Publicado: (2024)
Reward Model Overoptimisation in Iterated RLHF
por: Wolf, Lorenz, et al.
Publicado: (2025)
por: Wolf, Lorenz, et al.
Publicado: (2025)
State-free Reinforcement Learning
por: Chen, Mingyu, et al.
Publicado: (2024)
por: Chen, Mingyu, et al.
Publicado: (2024)
Learning Rate-Free Reinforcement Learning: A Case for Model Selection with Non-Stationary Objectives
por: Afshar, Aida, et al.
Publicado: (2024)
por: Afshar, Aida, et al.
Publicado: (2024)
Compute Aligned Training: Optimizing for Test Time Inference
por: Ousherovitch, Adam, et al.
Publicado: (2026)
por: Ousherovitch, Adam, et al.
Publicado: (2026)
Improved Training Mechanism for Reinforcement Learning via Online Model Selection
por: Afshar, Aida, et al.
Publicado: (2025)
por: Afshar, Aida, et al.
Publicado: (2025)
How to Explore with Belief: State Entropy Maximization in POMDPs
por: Zamboni, Riccardo, et al.
Publicado: (2024)
por: Zamboni, Riccardo, et al.
Publicado: (2024)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
por: Zhang, Chen Bo Calvin, et al.
Publicado: (2024)
por: Zhang, Chen Bo Calvin, et al.
Publicado: (2024)
Information-Theoretic Reward Decomposition for Generalizable RLHF
por: Mao, Liyuan, et al.
Publicado: (2025)
por: Mao, Liyuan, et al.
Publicado: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
por: Miao, Yuchun, et al.
Publicado: (2024)
por: Miao, Yuchun, et al.
Publicado: (2024)
Towards Principled Unsupervised Multi-Agent Reinforcement Learning
por: Zamboni, Riccardo, et al.
Publicado: (2025)
por: Zamboni, Riccardo, et al.
Publicado: (2025)
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
por: Misra, Dipendra, et al.
Publicado: (2026)
por: Misra, Dipendra, et al.
Publicado: (2026)
Data-Driven Online Model Selection With Regret Guarantees
por: Pacchiano, Aldo, et al.
Publicado: (2023)
por: Pacchiano, Aldo, et al.
Publicado: (2023)
In-Context Learning for Pure Exploration
por: Russo, Alessio, et al.
Publicado: (2025)
por: Russo, Alessio, et al.
Publicado: (2025)
Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching
por: Fraschini, Andrea, et al.
Publicado: (2026)
por: Fraschini, Andrea, et al.
Publicado: (2026)
Multiple-policy Evaluation via Density Estimation
por: Chen, Yilei, et al.
Publicado: (2024)
por: Chen, Yilei, et al.
Publicado: (2024)
Experiment Planning with Function Approximation
por: Pacchiano, Aldo, et al.
Publicado: (2024)
por: Pacchiano, Aldo, et al.
Publicado: (2024)
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
por: Tenedini, Davide, et al.
Publicado: (2025)
por: Tenedini, Davide, et al.
Publicado: (2025)
Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF
por: Liu, Jing
Publicado: (2025)
por: Liu, Jing
Publicado: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
por: Dong, Hanze, et al.
Publicado: (2024)
por: Dong, Hanze, et al.
Publicado: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Reward-Robust RLHF in LLMs
por: Yan, Yuzi, et al.
Publicado: (2024)
por: Yan, Yuzi, et al.
Publicado: (2024)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
Reward Compatibility: A Framework for Inverse RL
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
por: Zhang, Chuheng, et al.
Publicado: (2024)
por: Zhang, Chuheng, et al.
Publicado: (2024)
In-Context Learning for Pure Exploration in Continuous Spaces
por: Russo, Alessio, et al.
Publicado: (2026)
por: Russo, Alessio, et al.
Publicado: (2026)
Provable Interactive Learning with Hindsight Instruction Feedback
por: Misra, Dipendra, et al.
Publicado: (2024)
por: Misra, Dipendra, et al.
Publicado: (2024)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
por: Baharav, Tavor Z., et al.
Publicado: (2025)
por: Baharav, Tavor Z., et al.
Publicado: (2025)
How to Evaluate Reward Models for RLHF
por: Frick, Evan, et al.
Publicado: (2024)
por: Frick, Evan, et al.
Publicado: (2024)
A Characterization of List Language Identification in the Limit
por: Charikar, Moses, et al.
Publicado: (2025)
por: Charikar, Moses, et al.
Publicado: (2025)
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
por: Tang, Kenton, et al.
Publicado: (2026)
por: Tang, Kenton, et al.
Publicado: (2026)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
por: Zhao, Kangwen, et al.
Publicado: (2025)
por: Zhao, Kangwen, et al.
Publicado: (2025)
Quantile Regression for Distributional Reward Models in RLHF
por: Dorka, Nicolai
Publicado: (2024)
por: Dorka, Nicolai
Publicado: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
Accelerating RLHF Training with Reward Variance Increase
por: Yang, Zonglin, et al.
Publicado: (2025)
por: Yang, Zonglin, et al.
Publicado: (2025)
Ejemplares similares
-
Leveraging Offline Data in Linear Latent Contextual Bandits
por: Kausik, Chinmaya, et al.
Publicado: (2024) -
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
por: Hong, Kihyuk, et al.
Publicado: (2025) -
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
por: Kausik, Chinmaya, et al.
Publicado: (2026) -
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
por: Russo, Alessio, et al.
Publicado: (2025) -
Second Order Bounds for Contextual Bandits with Function Approximation
por: Pacchiano, Aldo
Publicado: (2024)