Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Shenao, Yu, Donghan, Sharma, Hiteshi, Zhong, Han, Liu, Zhihan, Yang, Ziyi, Wang, Shuohang, Hassan, Hany, Wang, Zhaoran |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning
by: Chen, Yifang, et al.
Published: (2024)
by: Chen, Yifang, et al.
Published: (2024)
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following
by: Yang, Yuxiao, et al.
Published: (2024)
by: Yang, Yuxiao, et al.
Published: (2024)
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
by: Xu, Ruijie, et al.
Published: (2024)
by: Xu, Ruijie, et al.
Published: (2024)
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Temperature-Centric Investigation of Speculative Decoding with Knowledge Distillation
by: Ouyang, Siru, et al.
Published: (2024)
by: Ouyang, Siru, et al.
Published: (2024)
Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample Efficiency
by: Liu, Zhihan, et al.
Published: (2023)
by: Liu, Zhihan, et al.
Published: (2023)
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
by: Zhong, Han, et al.
Published: (2025)
by: Zhong, Han, et al.
Published: (2025)
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
Online Self-Preferring Language Models
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
How Can LLM Guide RL? A Value-Based Approach
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Multi-LoRA Composition for Image Generation
by: Zhong, Ming, et al.
Published: (2024)
by: Zhong, Ming, et al.
Published: (2024)
Bayesian Preference Elicitation with Language Models
by: Handa, Kunal, et al.
Published: (2024)
by: Handa, Kunal, et al.
Published: (2024)
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
by: Bose, Avinandan, et al.
Published: (2024)
by: Bose, Avinandan, et al.
Published: (2024)
SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment
by: Zhu, Wenqiao, et al.
Published: (2025)
by: Zhu, Wenqiao, et al.
Published: (2025)
Contrastive UCB: Provably Efficient Contrastive Self-Supervised Learning in Online Reinforcement Learning
by: Qiu, Shuang, et al.
Published: (2022)
by: Qiu, Shuang, et al.
Published: (2022)
Self-Play Preference Optimization for Language Model Alignment
by: Wu, Yue, et al.
Published: (2024)
by: Wu, Yue, et al.
Published: (2024)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
by: Zhou, Han, et al.
Published: (2024)
by: Zhou, Han, et al.
Published: (2024)
On the Pros and Cons of Active Learning for Moral Preference Elicitation
by: Keswani, Vijay, et al.
Published: (2024)
by: Keswani, Vijay, et al.
Published: (2024)
Human-Instruction-Free LLM Self-Alignment with Limited Samples
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability
by: Hu, Jiachen, et al.
Published: (2025)
by: Hu, Jiachen, et al.
Published: (2025)
Online Preference Alignment for Language Models via Count-based Exploration
by: Bai, Chenjia, et al.
Published: (2025)
by: Bai, Chenjia, et al.
Published: (2025)
Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?
by: Gu, Zhuojun, et al.
Published: (2025)
by: Gu, Zhuojun, et al.
Published: (2025)
Human Alignment of Large Language Models through Online Preference Optimisation
by: Calandriello, Daniele, et al.
Published: (2024)
by: Calandriello, Daniele, et al.
Published: (2024)
Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis
by: Taddei, Sophia, et al.
Published: (2026)
by: Taddei, Sophia, et al.
Published: (2026)
SciAgent: Tool-augmented Language Models for Scientific Reasoning
by: Ma, Yubo, et al.
Published: (2024)
by: Ma, Yubo, et al.
Published: (2024)
Self-Augmented Preference Optimization: Off-Policy Paradigms for Language Model Alignment
by: Yin, Yueqin, et al.
Published: (2024)
by: Yin, Yueqin, et al.
Published: (2024)
Exploring Large Language Models for Climate Forecasting
by: Wang, Yang, et al.
Published: (2024)
by: Wang, Yang, et al.
Published: (2024)
Asking Clarifying Questions for Preference Elicitation With Large Language Models
by: Montazeralghaem, Ali, et al.
Published: (2025)
by: Montazeralghaem, Ali, et al.
Published: (2025)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
Large Language Model Supply Chain: A Research Agenda
by: Wang, Shenao, et al.
Published: (2024)
by: Wang, Shenao, et al.
Published: (2024)
Progressive Safeguards for Safe and Model-Agnostic Reinforcement Learning
by: Omi, Nabil, et al.
Published: (2024)
by: Omi, Nabil, et al.
Published: (2024)
Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment
by: Yang, Rui, et al.
Published: (2024)
by: Yang, Rui, et al.
Published: (2024)
SoK: Understanding Vulnerabilities in the Large Language Model Supply Chain
by: Wang, Shenao, et al.
Published: (2025)
by: Wang, Shenao, et al.
Published: (2025)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
Instant Preference Alignment for Text-to-Image Diffusion Models
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Eliciting Informed Preferences
by: Camara, Modibo K., et al.
Published: (2025)
by: Camara, Modibo K., et al.
Published: (2025)
Similar Items
-
Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning
by: Chen, Yifang, et al.
Published: (2024) -
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
by: Liu, Zhihan, et al.
Published: (2024) -
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024) -
Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following
by: Yang, Yuxiao, et al.
Published: (2024) -
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
by: Xu, Ruijie, et al.
Published: (2024)