Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Ziyu, Agarwal, Rishabh, Liu, Tianqi, Joshi, Rishabh, Velury, Sarmishta, Le, Quoc V., Tan, Qijun, Liu, Yuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BlendFusion -- Scalable Synthetic Data Generation for Diffusion Model Training
by: Venkatesh, Thejas, et al.
Published: (2026)
by: Venkatesh, Thejas, et al.
Published: (2026)
SiT: Symmetry-Invariant Transformers for Generalisation in Reinforcement Learning
by: Weissenbacher, Matthias, et al.
Published: (2024)
by: Weissenbacher, Matthias, et al.
Published: (2024)
From Camera to World: A Plug-and-Play Module for Human Mesh Transformation
by: Ma, Changhai, et al.
Published: (2025)
by: Ma, Changhai, et al.
Published: (2025)
Human Alignment of Large Language Models through Online Preference Optimisation
by: Calandriello, Daniele, et al.
Published: (2024)
by: Calandriello, Daniele, et al.
Published: (2024)
Enhancing Cache-Augmented Generation (CAG) with Adaptive Contextual Compression for Scalable Knowledge Integration
by: Agrawal, Rishabh, et al.
Published: (2025)
by: Agrawal, Rishabh, et al.
Published: (2025)
Statistical Rejection Sampling Improves Preference Optimization
by: Liu, Tianqi, et al.
Published: (2023)
by: Liu, Tianqi, et al.
Published: (2023)
Unified Knowledge Distillation Framework: Fine-Grained Alignment and Geometric Relationship Preservation for Deep Face Recognition
by: Mishra, Durgesh, et al.
Published: (2025)
by: Mishra, Durgesh, et al.
Published: (2025)
Crafting Tomorrow: The Influence of Design Choices on Fresh Content in Social Media Recommendation
by: Saket, Srijan, et al.
Published: (2024)
by: Saket, Srijan, et al.
Published: (2024)
The Art of Scaling Reinforcement Learning Compute for LLMs
by: Khatri, Devvrit, et al.
Published: (2025)
by: Khatri, Devvrit, et al.
Published: (2025)
Effect of Ethoxy functionalized hexagonal boron nitride on h‐ BN / TPU thermoplastic polyurethane nanocomposite and its thermal properties
by: Rishabh Tiwari, et al.
Published: (2024)
by: Rishabh Tiwari, et al.
Published: (2024)
A CRITICAL ANALYSIS OF FACILITATIVE MEDIATION AS A METHOD OF MEDIATION
by: Jain, Rishabh
Published: (2025)
by: Jain, Rishabh
Published: (2025)
A CRITICAL ASSESSMENT OF THE OCCUPATIONAL SAFETY, HEALTH AND WORKING CONDITIONS CODE AND ITS IMPLICATIONS FOR INFORMAL WORKERS
by: Rishabh Verma
Published: (2026)
by: Rishabh Verma
Published: (2026)
Spectral Reconstruction for Under-Resolved Turbulence Measurements Using a Variational Cutoff Dissipation Model
by: Mishra, Rishabh
Published: (2025)
by: Mishra, Rishabh
Published: (2025)
$c=1$, $R=1$ and $N\gg 1$: ZZ instantons in 2D String Theory and Matrix Integrals
by: Kaushik, Rishabh
Published: (2025)
by: Kaushik, Rishabh
Published: (2025)
Introduction to Sachdev-Ye-Kitaev Model: A Strongly Correlated System Perspective
by: Jha, Rishabh
Published: (2025)
by: Jha, Rishabh
Published: (2025)
Using text embedding models as text classifiers with medical data
by: Goel, Rishabh
Published: (2024)
by: Goel, Rishabh
Published: (2024)
Global Well-Posedness for the 3D Navier-Stokes Equations under Logarithmically Improved Criteria: Connections to Turbulence Theory
by: Mishra, Rishabh
Published: (2025)
by: Mishra, Rishabh
Published: (2025)
Runtime Evaluation of Procedural Content Generation in an Endless Runner Game Using Autonomous Agents
by: Kar, Rishabh
Published: (2026)
by: Kar, Rishabh
Published: (2026)
Adaptive Few-Shot Learning (AFSL): Tackling Data Scarcity with Stability, Robustness, and Versatility
by: Agrawal, Rishabh
Published: (2025)
by: Agrawal, Rishabh
Published: (2025)
Global Well-Posedness of the 3D Navier-Stokes Equations in the Limiting Case: Infinitely Nested Logarithmic Improvements
by: Mishra, Rishabh
Published: (2025)
by: Mishra, Rishabh
Published: (2025)
Talent allocation, gender disparities and post‐reform economic growth in Central America
by: Rishabh Sinha
Published: (2025)
by: Rishabh Sinha
Published: (2025)
The role of the second normal stress difference in rod-climbing effect
by: More, Rishabh
Published: (2025)
by: More, Rishabh
Published: (2025)
Global Well-Posedness of the 3D Navier-Stokes Equations under Multi-Level Logarithmically Improved Criteria
by: Mishra, Rishabh
Published: (2025)
by: Mishra, Rishabh
Published: (2025)
Complement Submodular Information Measures for Balanced and Robust Data Selection
by: Iyer, Rishabh
Published: (2026)
by: Iyer, Rishabh
Published: (2026)
Locked Subharmonic Oscillations in the Entanglement Spectrum of a Periodically Driven Topological Chain
by: Jha, Rishabh
Published: (2026)
by: Jha, Rishabh
Published: (2026)
Explicit Entropic Constructions for Coverage, Facility Location, and Graph Cuts
by: Iyer, Rishabh
Published: (2026)
by: Iyer, Rishabh
Published: (2026)
Radiatively Cooled Magnetic Reconnection Experiments Driven by Pulsed Power
by: Datta, Rishabh
Published: (2024)
by: Datta, Rishabh
Published: (2024)
Transformation and symmetries for the Andrews-Garvan crank function
by: Sarma, Rishabh
Published: (2023)
by: Sarma, Rishabh
Published: (2023)
How do servant leaders ignite absorptive capacity? The role of epistemic motivation and organizational support
by: Rishabh Rai
Published: (2016)
by: Rishabh Rai
Published: (2016)
SecurePose: Automated Face Blurring and Human Movement Kinematics Extraction from Videos Recorded in Clinical Settings
by: Bajpai, Rishabh, et al.
Published: (2024)
by: Bajpai, Rishabh, et al.
Published: (2024)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
V-STaR: Training Verifiers for Self-Taught Reasoners
by: Hosseini, Arian, et al.
Published: (2024)
by: Hosseini, Arian, et al.
Published: (2024)
Offline Regularised Reinforcement Learning for Large Language Models Alignment
by: Richemond, Pierre Harvey, et al.
Published: (2024)
by: Richemond, Pierre Harvey, et al.
Published: (2024)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
by: Sun, Zhiqing, et al.
Published: (2024)
by: Sun, Zhiqing, et al.
Published: (2024)
Training-Free Cross-Architecture Merging for Graph Neural Networks
by: Bhattacharya, Rishabh, et al.
Published: (2026)
by: Bhattacharya, Rishabh, et al.
Published: (2026)
Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
by: Majee, Anay, et al.
Published: (2025)
by: Majee, Anay, et al.
Published: (2025)
Coverage
by: Ravi Kumar, Rishabh
Published: (2025)
by: Ravi Kumar, Rishabh
Published: (2025)
Strain-transport superposition in shear-thinning dense non-Brownian suspensions
by: More, Rishabh V.
Published: (2026)
by: More, Rishabh V.
Published: (2026)
Adaptive Sliding Mode Control for Vehicle Platoons with State-Dependent Friction Uncertainty
by: Yadav, Rishabh Dev
Published: (2025)
by: Yadav, Rishabh Dev
Published: (2025)
Indigenous peoples in the world of work in Asia and the Pacific: a status report
by: Rishabh Kumar Dhir
Published: (2015)
by: Rishabh Kumar Dhir
Published: (2015)
Similar Items
-
BlendFusion -- Scalable Synthetic Data Generation for Diffusion Model Training
by: Venkatesh, Thejas, et al.
Published: (2026) -
SiT: Symmetry-Invariant Transformers for Generalisation in Reinforcement Learning
by: Weissenbacher, Matthias, et al.
Published: (2024) -
From Camera to World: A Plug-and-Play Module for Human Mesh Transformation
by: Ma, Changhai, et al.
Published: (2025) -
Human Alignment of Large Language Models through Online Preference Optimisation
by: Calandriello, Daniele, et al.
Published: (2024) -
Enhancing Cache-Augmented Generation (CAG) with Adaptive Contextual Compression for Scalable Knowledge Integration
by: Agrawal, Rishabh, et al.
Published: (2025)