Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling
Fuente:
arXiv
Guardado en:
| Autores principales: | Qi, Han, Yang, Haochen, Zhang, Qiaosheng, Yang, Zhuoran |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Provably Efficient Information-Directed Sampling Algorithms for Multi-Agent Reinforcement Learning
por: Zhang, Qiaosheng, et al.
Publicado: (2024)
por: Zhang, Qiaosheng, et al.
Publicado: (2024)
Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample Efficiency
por: Cai, Qi, et al.
Publicado: (2022)
por: Cai, Qi, et al.
Publicado: (2022)
Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation
por: He, Jianliang, et al.
Publicado: (2024)
por: He, Jianliang, et al.
Publicado: (2024)
Graph Feedback Bandits on Similar Arms: With and Without Graph Structures
por: Qi, Han, et al.
Publicado: (2025)
por: Qi, Han, et al.
Publicado: (2025)
Embed to Control Partially Observed Systems: Representation Learning with Provable Sample Efficiency
por: Wang, Lingxiao, et al.
Publicado: (2022)
por: Wang, Lingxiao, et al.
Publicado: (2022)
Sample-Efficient Policy Constraint Offline Deep Reinforcement Learning based on Sample Filtering
por: Chen, Yuanhao, et al.
Publicado: (2025)
por: Chen, Yuanhao, et al.
Publicado: (2025)
Sample Efficient Reinforcement Learning by Automatically Learning to Compose Subtasks
por: Han, Shuai, et al.
Publicado: (2024)
por: Han, Shuai, et al.
Publicado: (2024)
On the Role of Information Structure in Reinforcement Learning for Partially-Observable Sequential Teams and Games
por: Altabaa, Awni, et al.
Publicado: (2024)
por: Altabaa, Awni, et al.
Publicado: (2024)
GCHR : Goal-Conditioned Hindsight Regularization for Sample-Efficient Reinforcement Learning
por: Lei, Xing, et al.
Publicado: (2025)
por: Lei, Xing, et al.
Publicado: (2025)
More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling
por: Ishfaq, Haque, et al.
Publicado: (2024)
por: Ishfaq, Haque, et al.
Publicado: (2024)
Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
por: Shen, Han, et al.
Publicado: (2024)
por: Shen, Han, et al.
Publicado: (2024)
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
por: Zhang, Kongcheng, et al.
Publicado: (2025)
por: Zhang, Kongcheng, et al.
Publicado: (2025)
The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability
por: Hu, Jiachen, et al.
Publicado: (2025)
por: Hu, Jiachen, et al.
Publicado: (2025)
Provably Sample-Efficient Robust Reinforcement Learning with Average Reward
por: Roch, Zachary, et al.
Publicado: (2025)
por: Roch, Zachary, et al.
Publicado: (2025)
Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference
por: Cercola, Matteo, et al.
Publicado: (2025)
por: Cercola, Matteo, et al.
Publicado: (2025)
Optimistic Information Directed Sampling
por: Neu, Gergely, et al.
Publicado: (2024)
por: Neu, Gergely, et al.
Publicado: (2024)
Adaptive Client Sampling in Federated Learning via Online Learning with Bandit Feedback
por: Zhao, Boxin, et al.
Publicado: (2021)
por: Zhao, Boxin, et al.
Publicado: (2021)
Contrastive UCB: Provably Efficient Contrastive Self-Supervised Learning in Online Reinforcement Learning
por: Qiu, Shuang, et al.
Publicado: (2022)
por: Qiu, Shuang, et al.
Publicado: (2022)
Reinforcement Learning from Human Feedback
por: Lambert, Nathan
Publicado: (2025)
por: Lambert, Nathan
Publicado: (2025)
Parameter Efficient Reinforcement Learning from Human Feedback
por: Sidahmed, Hakim, et al.
Publicado: (2024)
por: Sidahmed, Hakim, et al.
Publicado: (2024)
SEAR: Sample Efficient Action Chunking Reinforcement Learning
por: Nagy, C. F. Maximilian, et al.
Publicado: (2026)
por: Nagy, C. F. Maximilian, et al.
Publicado: (2026)
Sample Efficient Reinforcement Learning via Large Vision Language Model Distillation
por: Lee, Donghoon, et al.
Publicado: (2025)
por: Lee, Donghoon, et al.
Publicado: (2025)
Robust Reinforcement Learning from Corrupted Human Feedback
por: Bukharin, Alexander, et al.
Publicado: (2024)
por: Bukharin, Alexander, et al.
Publicado: (2024)
Sample-Efficient Tabular Self-Play for Offline Robust Reinforcement Learning
por: Li, Na, et al.
Publicado: (2025)
por: Li, Na, et al.
Publicado: (2025)
On Sample-Efficient Offline Reinforcement Learning: Data Diversity, Posterior Sampling, and Beyond
por: Nguyen-Tang, Thanh, et al.
Publicado: (2024)
por: Nguyen-Tang, Thanh, et al.
Publicado: (2024)
Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation
por: Zhao, Runze, et al.
Publicado: (2025)
por: Zhao, Runze, et al.
Publicado: (2025)
TIC-GRPO: Provable and Efficient Optimization for Reinforcement Learning from Human Feedback
por: Pang, Lei, et al.
Publicado: (2025)
por: Pang, Lei, et al.
Publicado: (2025)
Sparse Optimistic Information Directed Sampling
por: Schwartz, Ludovic, et al.
Publicado: (2025)
por: Schwartz, Ludovic, et al.
Publicado: (2025)
From Generative to Episodic: Sample-Efficient Replicable Reinforcement Learning
por: Hopkins, Max, et al.
Publicado: (2025)
por: Hopkins, Max, et al.
Publicado: (2025)
Strategyproof Reinforcement Learning from Human Feedback
por: Buening, Thomas Kleine, et al.
Publicado: (2025)
por: Buening, Thomas Kleine, et al.
Publicado: (2025)
CoPRIS: Efficient and Stable Reinforcement Learning via Concurrency-Controlled Partial Rollout with Importance Sampling
por: Qu, Zekai, et al.
Publicado: (2025)
por: Qu, Zekai, et al.
Publicado: (2025)
Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning
por: Liu, Xu-Hui, et al.
Publicado: (2024)
por: Liu, Xu-Hui, et al.
Publicado: (2024)
Community Detection for Contextual-LSBM: Theoretical Limitations of Misclassification Rate and Efficient Algorithms
por: Jin, Dian, et al.
Publicado: (2025)
por: Jin, Dian, et al.
Publicado: (2025)
Truncated Rectified Flow Policy for Reinforcement Learning with One-Step Sampling
por: Zhou, Xubin, et al.
Publicado: (2026)
por: Zhou, Xubin, et al.
Publicado: (2026)
Aligning AI Agents via Information-Directed Sampling
por: Jeon, Hong Jun, et al.
Publicado: (2024)
por: Jeon, Hong Jun, et al.
Publicado: (2024)
Quantum Boltzmann Machines for Sample-Efficient Reinforcement Learning
por: Gerlach, Thore, et al.
Publicado: (2025)
por: Gerlach, Thore, et al.
Publicado: (2025)
Sample-Efficient Reinforcement Learning of Koopman eNMPC
por: Mayfrank, Daniel, et al.
Publicado: (2025)
por: Mayfrank, Daniel, et al.
Publicado: (2025)
Sample-Efficient Constrained Reinforcement Learning with General Parameterization
por: Mondal, Washim Uddin, et al.
Publicado: (2024)
por: Mondal, Washim Uddin, et al.
Publicado: (2024)
Sample Efficient Reinforcement Learning with Partial Dynamics Knowledge
por: Alharbi, Meshal, et al.
Publicado: (2023)
por: Alharbi, Meshal, et al.
Publicado: (2023)
Sample Efficient Active Algorithms for Offline Reinforcement Learning
por: Roy, Soumyadeep, et al.
Publicado: (2026)
por: Roy, Soumyadeep, et al.
Publicado: (2026)
Ejemplares similares
-
Provably Efficient Information-Directed Sampling Algorithms for Multi-Agent Reinforcement Learning
por: Zhang, Qiaosheng, et al.
Publicado: (2024) -
Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample Efficiency
por: Cai, Qi, et al.
Publicado: (2022) -
Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation
por: He, Jianliang, et al.
Publicado: (2024) -
Graph Feedback Bandits on Similar Arms: With and Without Graph Structures
por: Qi, Han, et al.
Publicado: (2025) -
Embed to Control Partially Observed Systems: Representation Learning with Provable Sample Efficiency
por: Wang, Lingxiao, et al.
Publicado: (2022)