Best Policy Learning from Trajectory Preference Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Agnihotri, Akhil, Jain, Rahul, Ramachandran, Deepak, Wen, Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Online Bandit Learning with Offline Preference Data for Improved RLHF
by: Agnihotri, Akhil, et al.
Published: (2024)
by: Agnihotri, Akhil, et al.
Published: (2024)
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
by: Agnihotri, Akhil, et al.
Published: (2025)
by: Agnihotri, Akhil, et al.
Published: (2025)
e-COP : Episodic Constrained Optimization of Policies
by: Agnihotri, Akhil, et al.
Published: (2024)
by: Agnihotri, Akhil, et al.
Published: (2024)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
Multi-Objective Reward and Preference Optimization: Theory and Algorithms
by: Agnihotri, Akhil
Published: (2025)
by: Agnihotri, Akhil
Published: (2025)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
by: Xu, Zaiyan, et al.
Published: (2025)
by: Xu, Zaiyan, et al.
Published: (2025)
Best-of-Both-Worlds Policy Optimization for CMDPs with Bandit Feedback
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
RANDPOL: Parameter-Efficient End-to-End Quadruped Locomotion via Randomized Policy Learning
by: Liu, Zhuochen, et al.
Published: (2025)
by: Liu, Zhuochen, et al.
Published: (2025)
Pure Exploration for Constrained Best Mixed Arm Identification with a Fixed Budget
by: Tang, Dengwang, et al.
Published: (2024)
by: Tang, Dengwang, et al.
Published: (2024)
Off-Policy Evaluation from Logged Human Feedback
by: Bhargava, Aniruddha, et al.
Published: (2024)
by: Bhargava, Aniruddha, et al.
Published: (2024)
Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
by: Alwis, Praditha, et al.
Published: (2026)
by: Alwis, Praditha, et al.
Published: (2026)
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2025)
by: Burnwal, Returaj, et al.
Published: (2025)
Combinatorial Reinforcement Learning with Preference Feedback
by: Lee, Joongkyu, et al.
Published: (2025)
by: Lee, Joongkyu, et al.
Published: (2025)
Selection of the Best Policy under Fairness Constraints for Subpopulations
by: Zhu, Tingyu, et al.
Published: (2026)
by: Zhu, Tingyu, et al.
Published: (2026)
Pref-GUIDE: Continual Policy Learning from Real-Time Human Feedback via Preference-Based Learning
by: Ji, Zhengran, et al.
Published: (2025)
by: Ji, Zhengran, et al.
Published: (2025)
Multi-turn Reinforcement Learning from Preference Human Feedback
by: Shani, Lior, et al.
Published: (2024)
by: Shani, Lior, et al.
Published: (2024)
Online Policy Learning from Offline Preferences
by: Zhang, Guoxi, et al.
Published: (2024)
by: Zhang, Guoxi, et al.
Published: (2024)
Efficient Online Learning with Offline Datasets for Infinite Horizon MDPs: A Bayesian Approach
by: Tang, Dengwang, et al.
Published: (2023)
by: Tang, Dengwang, et al.
Published: (2023)
DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)
by: Hejna, Joey, et al.
Published: (2023)
Online Iterative Reinforcement Learning from Human Feedback with General Preference Model
by: Ye, Chenlu, et al.
Published: (2024)
by: Ye, Chenlu, et al.
Published: (2024)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)
by: Cho, Taehyun, et al.
Published: (2025)
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles
by: Pukdee, Rattana, et al.
Published: (2026)
by: Pukdee, Rattana, et al.
Published: (2026)
Coresets from Trajectories: Selecting Data via Correlation of Loss Differences
by: Nagaraj, Manish, et al.
Published: (2025)
by: Nagaraj, Manish, et al.
Published: (2025)
Queueing Matching Bandits with Preference Feedback
by: Kim, Jung-hun, et al.
Published: (2024)
by: Kim, Jung-hun, et al.
Published: (2024)
Learning from Language Feedback via Variational Policy Distillation
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
by: Swamy, Gokul, et al.
Published: (2024)
by: Swamy, Gokul, et al.
Published: (2024)
Adaptive Preference Scaling for Reinforcement Learning with Human Feedback
by: Hong, Ilgee, et al.
Published: (2024)
by: Hong, Ilgee, et al.
Published: (2024)
Contextual Online Uncertainty-Aware Preference Learning for Human Feedback
by: Lu, Nan, et al.
Published: (2025)
by: Lu, Nan, et al.
Published: (2025)
Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback
by: Kim, Gihoon, et al.
Published: (2026)
by: Kim, Gihoon, et al.
Published: (2026)
Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
by: Schlisselberg, Ofir, et al.
Published: (2025)
by: Schlisselberg, Ofir, et al.
Published: (2025)
Query-Policy Misalignment in Preference-Based Reinforcement Learning
by: Hu, Xiao, et al.
Published: (2023)
by: Hu, Xiao, et al.
Published: (2023)
Preference-based Conditional Treatment Effects and Policy Learning
by: Parnas, Dovid, et al.
Published: (2026)
by: Parnas, Dovid, et al.
Published: (2026)
PTrajM: Efficient and Semantic-rich Trajectory Learning with Pretrained Trajectory-Mamba
by: Lin, Yan, et al.
Published: (2024)
by: Lin, Yan, et al.
Published: (2024)
Policy Teaching via Data Poisoning in Learning from Human Preferences
by: Nika, Andi, et al.
Published: (2025)
by: Nika, Andi, et al.
Published: (2025)
Are Retrials All You Need? Enhancing Large Language Model Reasoning Without Verbalized Feedback
by: Potamitis, Nearchos, et al.
Published: (2025)
by: Potamitis, Nearchos, et al.
Published: (2025)
Making RL with Preference-based Feedback Efficient via Randomization
by: Wu, Runzhe, et al.
Published: (2023)
by: Wu, Runzhe, et al.
Published: (2023)
Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
by: Zhang, Rongtao, et al.
Published: (2026)
by: Zhang, Rongtao, et al.
Published: (2026)
Multi-Phase Spacecraft Trajectory Optimization via Transformer-Based Reinforcement Learning
by: Jain, Amit, et al.
Published: (2025)
by: Jain, Amit, et al.
Published: (2025)
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance
by: Agnihotri, Rudransh, et al.
Published: (2025)
by: Agnihotri, Rudransh, et al.
Published: (2025)
Similar Items
-
Online Bandit Learning with Offline Preference Data for Improved RLHF
by: Agnihotri, Akhil, et al.
Published: (2024) -
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
by: Agnihotri, Akhil, et al.
Published: (2025) -
e-COP : Episodic Constrained Optimization of Policies
by: Agnihotri, Akhil, et al.
Published: (2024) -
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023) -
Multi-Objective Reward and Preference Optimization: Theory and Algorithms
by: Agnihotri, Akhil
Published: (2025)