Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Ruijie, Liu, Zhihan, Liu, Yongfei, Yan, Shipeng, Wang, Zhaoran, Zhang, Zhi, He, Xuming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
Vision Language Models See What You Want but not What You See
by: Gao, Qingying, et al.
Published: (2024)
by: Gao, Qingying, et al.
Published: (2024)
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
by: Zhang, Mingxuan, et al.
Published: (2026)
by: Zhang, Mingxuan, et al.
Published: (2026)
MOSLIM:Align with diverse preferences in prompts through reward classification
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following
by: Yang, Yuxiao, et al.
Published: (2024)
by: Yang, Yuxiao, et al.
Published: (2024)
DA-DPO: Cost-efficient Difficulty-aware Preference Optimization for Reducing MLLM Hallucinations
by: Qiu, Longtian, et al.
Published: (2026)
by: Qiu, Longtian, et al.
Published: (2026)
Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization
by: Dong, Zhijin
Published: (2025)
by: Dong, Zhijin
Published: (2025)
You Only Forward Once: An Efficient Compositional Judging Paradigm
by: Zhang, Tianlong, et al.
Published: (2025)
by: Zhang, Tianlong, et al.
Published: (2025)
What should be observed for optimal reward in POMDPs?
by: Konsta, Alyzia-Maria, et al.
Published: (2024)
by: Konsta, Alyzia-Maria, et al.
Published: (2024)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
by: Liu, Shih-Yang, et al.
Published: (2026)
by: Liu, Shih-Yang, et al.
Published: (2026)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
by: Qiu, Longtian, et al.
Published: (2024)
by: Qiu, Longtian, et al.
Published: (2024)
Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample Efficiency
by: Liu, Zhihan, et al.
Published: (2023)
by: Liu, Zhihan, et al.
Published: (2023)
OET: Optimization-based prompt injection Evaluation Toolkit
by: Pan, Jinsheng, et al.
Published: (2025)
by: Pan, Jinsheng, et al.
Published: (2025)
What Is AI Safety? What Do We Want It to Be?
by: Harding, Jacqueline, et al.
Published: (2025)
by: Harding, Jacqueline, et al.
Published: (2025)
Count What You Want: Exemplar Identification and Few-shot Counting of Human Actions in the Wild
by: Huang, Yifeng, et al.
Published: (2023)
by: Huang, Yifeng, et al.
Published: (2023)
MCTS-EP: Empowering Embodied Planning with Online Preference Optimization
by: Xu, Hang, et al.
Published: (2025)
by: Xu, Hang, et al.
Published: (2025)
Self-rewarding correction for mathematical reasoning
by: Xiong, Wei, et al.
Published: (2025)
by: Xiong, Wei, et al.
Published: (2025)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
by: Xu, Tianyang, et al.
Published: (2024)
by: Xu, Tianyang, et al.
Published: (2024)
Only Send What You Need: Learning to Communicate Efficiently in Federated Multilingual Machine Translation
by: Chu, Yun-Wei, et al.
Published: (2024)
by: Chu, Yun-Wei, et al.
Published: (2024)
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
What Do AI-Generated Images Want?
by: Wasielewski, Amanda
Published: (2025)
by: Wasielewski, Amanda
Published: (2025)
Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents
by: Wang, Yufeng
Published: (2026)
by: Wang, Yufeng
Published: (2026)
Using Multi-modal Large Language Model to Boost Fireworks Algorithm's Ability in Settling Challenging Optimization Tasks
by: Cen, Shipeng, et al.
Published: (2025)
by: Cen, Shipeng, et al.
Published: (2025)
Unlearning of Knowledge Graph Embedding via Preference Optimization
by: Liu, Jiajun, et al.
Published: (2025)
by: Liu, Jiajun, et al.
Published: (2025)
How Culture Shapes What People Want From AI
by: Ge, Xiao, et al.
Published: (2024)
by: Ge, Xiao, et al.
Published: (2024)
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
by: Sirdeshmukh, Ved, et al.
Published: (2026)
by: Sirdeshmukh, Ved, et al.
Published: (2026)
Does the Model Say What the Data Says? A Simple Heuristic for Model Data Alignment
by: Salgado, Henry, et al.
Published: (2025)
by: Salgado, Henry, et al.
Published: (2025)
Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences
by: Ramírez, Guillem, et al.
Published: (2025)
by: Ramírez, Guillem, et al.
Published: (2025)
Self-Consistency Preference Optimization
by: Prasad, Archiki, et al.
Published: (2024)
by: Prasad, Archiki, et al.
Published: (2024)
Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
by: Wu, Yilin, et al.
Published: (2025)
by: Wu, Yilin, et al.
Published: (2025)
Streaming Looking Ahead with Token-level Self-reward
by: Zhang, Hongming, et al.
Published: (2025)
by: Zhang, Hongming, et al.
Published: (2025)
Self-supervised Preference Optimization: Enhance Your Language Model with Preference Degree Awareness
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
Expertise Is What We Want
by: Ashworth, Alan, et al.
Published: (2025)
by: Ashworth, Alan, et al.
Published: (2025)
POPI: Personalizing LLMs via Optimized Natural Language Preference Inference
by: Chen, Yizhuo, et al.
Published: (2025)
by: Chen, Yizhuo, et al.
Published: (2025)
Preference learning in shades of gray: Interpretable and bias-aware reward modeling for human preferences
by: Oprea, Simona-Vasilica, et al.
Published: (2026)
by: Oprea, Simona-Vasilica, et al.
Published: (2026)
TSO: Self-Training with Scaled Preference Optimization
by: Chen, Kaihui, et al.
Published: (2024)
by: Chen, Kaihui, et al.
Published: (2024)
Small-Margin Preferences Still Matter-If You Train Them Right
by: Pang, Jinlong, et al.
Published: (2026)
by: Pang, Jinlong, et al.
Published: (2026)
Similar Items
-
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
by: Liu, Zhihan, et al.
Published: (2024) -
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024) -
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
by: Zhang, Shenao, et al.
Published: (2024) -
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
by: Guo, Hongyi, et al.
Published: (2024) -
Vision Language Models See What You Want but not What You See
by: Gao, Qingying, et al.
Published: (2024)