Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ji, Xiang, Kulkarni, Sanjeev, Wang, Mengdi, Xie, Tengyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Play Preference Optimization for Language Model Alignment
von: Wu, Yue, et al.
Veröffentlicht: (2024)
von: Wu, Yue, et al.
Veröffentlicht: (2024)
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
Reinforce LLM Reasoning through Multi-Agent Reflection
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
von: Chen, Zixiang, et al.
Veröffentlicht: (2024)
von: Chen, Zixiang, et al.
Veröffentlicht: (2024)
A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement
von: Yuan, Hui, et al.
Veröffentlicht: (2024)
von: Yuan, Hui, et al.
Veröffentlicht: (2024)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
RoleCraft-GLM: Advancing Personalized Role-Playing in Large Language Models
von: Tao, Meiling, et al.
Veröffentlicht: (2023)
von: Tao, Meiling, et al.
Veröffentlicht: (2023)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
von: Wei, Boyi, et al.
Veröffentlicht: (2024)
von: Wei, Boyi, et al.
Veröffentlicht: (2024)
DeepCritic: Deliberate Critique with Large Language Models
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
von: Shen, Gerald, et al.
Veröffentlicht: (2024)
von: Shen, Gerald, et al.
Veröffentlicht: (2024)
Extensive Self-Contrast Enables Feedback-Free Language Model Alignment
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
Towards Scalable Automated Alignment of LLMs: A Survey
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
Offline Learning and Forgetting for Reasoning with Large Language Models
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
von: Chen, Yanxi, et al.
Veröffentlicht: (2024)
von: Chen, Yanxi, et al.
Veröffentlicht: (2024)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
Why is Your Language Model a Poor Implicit Reward Model?
von: Razin, Noam, et al.
Veröffentlicht: (2025)
von: Razin, Noam, et al.
Veröffentlicht: (2025)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)
Model Extrapolation Expedites Alignment
von: Zheng, Chujie, et al.
Veröffentlicht: (2024)
von: Zheng, Chujie, et al.
Veröffentlicht: (2024)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
von: Wang, Ru, et al.
Veröffentlicht: (2025)
von: Wang, Ru, et al.
Veröffentlicht: (2025)
CriticAL: Critic Automation with Language Models
von: Li, Michael Y., et al.
Veröffentlicht: (2024)
von: Li, Michael Y., et al.
Veröffentlicht: (2024)
Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play
von: Ye, Ziyu, et al.
Veröffentlicht: (2024)
von: Ye, Ziyu, et al.
Veröffentlicht: (2024)
SALMON: Self-Alignment with Instructable Reward Models
von: Sun, Zhiqing, et al.
Veröffentlicht: (2023)
von: Sun, Zhiqing, et al.
Veröffentlicht: (2023)
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
von: Li, Chengao, et al.
Veröffentlicht: (2025)
von: Li, Chengao, et al.
Veröffentlicht: (2025)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
von: Yuan, Huizhuo, et al.
Veröffentlicht: (2024)
von: Yuan, Huizhuo, et al.
Veröffentlicht: (2024)
Can Brain Signals Reveal Inner Alignment with Human Languages?
von: Han, William, et al.
Veröffentlicht: (2022)
von: Han, William, et al.
Veröffentlicht: (2022)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
Recall-Extend Dynamics: Enhancing Small Language Models through Controlled Exploration and Refined Offline Integration
von: Guan, Zhong, et al.
Veröffentlicht: (2025)
von: Guan, Zhong, et al.
Veröffentlicht: (2025)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
von: Patel, Dev, et al.
Veröffentlicht: (2025)
von: Patel, Dev, et al.
Veröffentlicht: (2025)
On the Robustness of Reward Models for Language Model Alignment
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
Adversarial Reinforcement Learning for Large Language Model Agent Safety
von: Wang, Zizhao, et al.
Veröffentlicht: (2025)
von: Wang, Zizhao, et al.
Veröffentlicht: (2025)
Probabilistic Token Alignment for Large Language Model Fusion
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Self-Play Preference Optimization for Language Model Alignment
von: Wu, Yue, et al.
Veröffentlicht: (2024) -
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
von: Rosset, Corby, et al.
Veröffentlicht: (2024) -
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025) -
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
von: Yuan, Yurun, et al.
Veröffentlicht: (2026) -
Reinforce LLM Reasoning through Multi-Agent Reflection
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)