Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Audrey, Block, Adam, Liu, Qinghua, Jiang, Nan, Krishnamurthy, Akshay, Foster, Dylan J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting the (Sub)Optimality of Best-of-N for Inference-Time Alignment
by: Sriraman, Ved, et al.
Published: (2026)
by: Sriraman, Ved, et al.
Published: (2026)
A Unifying View of Coverage in Linear Off-Policy Evaluation
by: Amortila, Philip, et al.
Published: (2026)
by: Amortila, Philip, et al.
Published: (2026)
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Self-Improvement in Language Models: The Sharpening Mechanism
by: Huang, Audrey, et al.
Published: (2024)
by: Huang, Audrey, et al.
Published: (2024)
Representation-Based Exploration for Language Models: From Test-Time to Post-Training
by: Tuyls, Jens, et al.
Published: (2025)
by: Tuyls, Jens, et al.
Published: (2025)
Best-of-Tails: Bridging Optimism and Pessimism in Inference-Time Alignment
by: Hsu, Hsiang, et al.
Published: (2026)
by: Hsu, Hsiang, et al.
Published: (2026)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
by: Huang, Audrey, et al.
Published: (2024)
by: Huang, Audrey, et al.
Published: (2024)
Reinforcement Learning under Latent Dynamics: Toward Statistical and Algorithmic Modularity
by: Amortila, Philip, et al.
Published: (2024)
by: Amortila, Philip, et al.
Published: (2024)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
by: Qiu, Jiahao, et al.
Published: (2024)
by: Qiu, Jiahao, et al.
Published: (2024)
Variational Best-of-N Alignment
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification
by: Rohatgi, Dhruv, et al.
Published: (2025)
by: Rohatgi, Dhruv, et al.
Published: (2025)
Is Behavior Cloning All You Need? Understanding Horizon in Imitation Learning
by: Foster, Dylan J., et al.
Published: (2024)
by: Foster, Dylan J., et al.
Published: (2024)
AdaBoN: Adaptive Best-of-N Alignment
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
Can large language models explore in-context?
by: Krishnamurthy, Akshay, et al.
Published: (2024)
by: Krishnamurthy, Akshay, et al.
Published: (2024)
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024)
by: Hughes, John, et al.
Published: (2024)
The Best Instruction-Tuning Data are Those That Fit
by: Zhang, Dylan, et al.
Published: (2025)
by: Zhang, Dylan, et al.
Published: (2025)
Majority of the Bests: Improving Best-of-N via Bootstrapping
by: Rakhsha, Amin, et al.
Published: (2025)
by: Rakhsha, Amin, et al.
Published: (2025)
Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference
by: Golowich, Noah, et al.
Published: (2026)
by: Golowich, Noah, et al.
Published: (2026)
Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
by: Rho, Hyung Gyu, et al.
Published: (2025)
by: Rho, Hyung Gyu, et al.
Published: (2025)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
by: Chow, Yinlam, et al.
Published: (2024)
by: Chow, Yinlam, et al.
Published: (2024)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
by: Xie, Tengyang, et al.
Published: (2024)
by: Xie, Tengyang, et al.
Published: (2024)
Learning Generative Selection for Best-of-N
by: Toshniwal, Shubham, et al.
Published: (2026)
by: Toshniwal, Shubham, et al.
Published: (2026)
Learning to Select the Best Forecasting Tasks for Clinical Outcome Prediction
by: Xue, Yuan, et al.
Published: (2024)
by: Xue, Yuan, et al.
Published: (2024)
Best-of-$\infty$ -- Asymptotic Performance of Test-Time LLM Ensembling
by: Komiyama, Junpei, et al.
Published: (2025)
by: Komiyama, Junpei, et al.
Published: (2025)
STEB: In Search of the Best Evaluation Approach for Synthetic Time Series
by: Stenger, Michael, et al.
Published: (2025)
by: Stenger, Michael, et al.
Published: (2025)
BOND: Aligning LLMs with Best-of-N Distillation
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
Identifying the Best Transition Law
by: Ahmadipour, Mehrasa, et al.
Published: (2025)
by: Ahmadipour, Mehrasa, et al.
Published: (2025)
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
by: Tang, Yung-Chen, et al.
Published: (2025)
by: Tang, Yung-Chen, et al.
Published: (2025)
Optimal Multi-Objective Best Arm Identification with Fixed Confidence
by: Chen, Zhirui, et al.
Published: (2025)
by: Chen, Zhirui, et al.
Published: (2025)
LLMs as High-Dimensional Nonlinear Autoregressive Models with Attention: Training, Alignment and Inference
by: Krishnamurthy, Vikram
Published: (2026)
by: Krishnamurthy, Vikram
Published: (2026)
Best-Arm Identification in Unimodal Bandits
by: Poiani, Riccardo, et al.
Published: (2024)
by: Poiani, Riccardo, et al.
Published: (2024)
The Role of Environment Access in Agnostic Reinforcement Learning
by: Krishnamurthy, Akshay, et al.
Published: (2025)
by: Krishnamurthy, Akshay, et al.
Published: (2025)
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles
by: Pukdee, Rattana, et al.
Published: (2026)
by: Pukdee, Rattana, et al.
Published: (2026)
Asymptotically Optimal Linear Best Feasible Arm Identification with Fixed Budget
by: Bian, Jie, et al.
Published: (2025)
by: Bian, Jie, et al.
Published: (2025)
Generalized Neyman Allocation for Locally Minimax Optimal Best-Arm Identification
by: Kato, Masahiro
Published: (2024)
by: Kato, Masahiro
Published: (2024)
Fair Best Arm Identification with Fixed Confidence
by: Russo, Alessio, et al.
Published: (2024)
by: Russo, Alessio, et al.
Published: (2024)
Constrained Best Arm Identification with Tests for Feasibility
by: Cai, Ting, et al.
Published: (2025)
by: Cai, Ting, et al.
Published: (2025)
Multi-Armed Bandits With Best-Action Queries
by: Bacchiocchi, Francesco, et al.
Published: (2026)
by: Bacchiocchi, Francesco, et al.
Published: (2026)
Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits
by: Hou, Yunlong, et al.
Published: (2024)
by: Hou, Yunlong, et al.
Published: (2024)
Similar Items
-
Revisiting the (Sub)Optimality of Best-of-N for Inference-Time Alignment
by: Sriraman, Ved, et al.
Published: (2026) -
A Unifying View of Coverage in Linear Off-Policy Evaluation
by: Amortila, Philip, et al.
Published: (2026) -
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025) -
Self-Improvement in Language Models: The Sharpening Mechanism
by: Huang, Audrey, et al.
Published: (2024) -
Representation-Based Exploration for Language Models: From Test-Time to Post-Training
by: Tuyls, Jens, et al.
Published: (2025)