Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Tong, Mei, Jincheng, Dai, Hanjun, Wen, Zixin, Cen, Shicong, Schuurmans, Dale, Chi, Yuejie, Dai, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Expectations: Learning with Stochastic Dominance Made Practical
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
Stochastic Gradient Succeeds for Bandits
by: Mei, Jincheng, et al.
Published: (2024)
by: Mei, Jincheng, et al.
Published: (2024)
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
by: Yang, Tong, et al.
Published: (2023)
by: Yang, Tong, et al.
Published: (2023)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Autoregressive Large Language Models are Computationally Universal
by: Schuurmans, Dale, et al.
Published: (2024)
by: Schuurmans, Dale, et al.
Published: (2024)
Spectral Ghost in Representation Learning: from Component Analysis to Self-Supervised Learning
by: Dai, Bo, et al.
Published: (2026)
by: Dai, Bo, et al.
Published: (2026)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
by: Dong, Harry, et al.
Published: (2025)
by: Dong, Harry, et al.
Published: (2025)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Spectral Representation-based Reinforcement Learning
by: Gao, Chenxiao, et al.
Published: (2025)
by: Gao, Chenxiao, et al.
Published: (2025)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
by: Zhang, Hongming, et al.
Published: (2023)
by: Zhang, Hongming, et al.
Published: (2023)
A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
by: Huang, Yu, et al.
Published: (2024)
by: Huang, Yu, et al.
Published: (2024)
UQE: A Query Engine for Unstructured Databases
by: Dai, Hanjun, et al.
Published: (2024)
by: Dai, Hanjun, et al.
Published: (2024)
Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
by: Che, Fengdi, et al.
Published: (2024)
by: Che, Fengdi, et al.
Published: (2024)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
by: Bohnet, Bernd, et al.
Published: (2024)
by: Bohnet, Bernd, et al.
Published: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Diffusion Controller: Framework, Algorithms and Parameterization
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Towards Faster Non-Asymptotic Convergence for Diffusion-Based Generative Models
by: Li, Gen, et al.
Published: (2023)
by: Li, Gen, et al.
Published: (2023)
Delightful Gradients Accelerate Corner Escape
by: Mei, Jincheng, et al.
Published: (2026)
by: Mei, Jincheng, et al.
Published: (2026)
Representation Learning via Non-Contrastive Mutual Information
by: Guo, Zhaohan Daniel, et al.
Published: (2025)
by: Guo, Zhaohan Daniel, et al.
Published: (2025)
AmorLIP: Efficient Language-Image Pretraining via Amortization
by: Sun, Haotian, et al.
Published: (2025)
by: Sun, Haotian, et al.
Published: (2025)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Plastic Learning with Deep Fourier Features
by: Lewandowski, Alex, et al.
Published: (2024)
by: Lewandowski, Alex, et al.
Published: (2024)
Large Language Models can Learn Rules
by: Zhu, Zhaocheng, et al.
Published: (2023)
by: Zhu, Zhaocheng, et al.
Published: (2023)
Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity
by: Shi, Laixi, et al.
Published: (2022)
by: Shi, Laixi, et al.
Published: (2022)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
Variational Best-of-N Alignment
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
The Sample-Communication Complexity Trade-off in Federated Q-Learning
by: Salgia, Sudeep, et al.
Published: (2024)
by: Salgia, Sudeep, et al.
Published: (2024)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
by: Li, Changhao, et al.
Published: (2024)
by: Li, Changhao, et al.
Published: (2024)
Toward Understanding In-context vs. In-weight Learning
by: Chan, Bryan, et al.
Published: (2024)
by: Chan, Bryan, et al.
Published: (2024)
Preference Optimization for Molecule Synthesis with Conditional Residual Energy-based Models
by: Liu, Songtao, et al.
Published: (2024)
by: Liu, Songtao, et al.
Published: (2024)
FasterSTS: A Faster Spatio-Temporal Synchronous Graph Convolutional Networks for Traffic flow Forecasting
by: Dai, Ben-Ao, et al.
Published: (2025)
by: Dai, Ben-Ao, et al.
Published: (2025)
Communication-Efficient Federated Optimization over Semi-Decentralized Networks
by: Wang, He, et al.
Published: (2023)
by: Wang, He, et al.
Published: (2023)
BOND: Aligning LLMs with Best-of-N Distillation
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
Similar Items
-
Beyond Expectations: Learning with Stochastic Dominance Made Practical
by: Cen, Shicong, et al.
Published: (2024) -
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
by: Cen, Shicong, et al.
Published: (2024) -
Stochastic Gradient Succeeds for Bandits
by: Mei, Jincheng, et al.
Published: (2024) -
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
by: Yang, Tong, et al.
Published: (2023) -
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
by: Mei, Jincheng, et al.
Published: (2025)