Best-of-Majority: Minimax-Optimal Strategy for Pass@$k$ Inference Scaling
Fuente:
arXiv
Guardado en:
| Autores principales: | Di, Qiwei, Ji, Kaixuan, Li, Xuheng, Zhao, Heyang, Gu, Quanquan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
por: Ji, Kaixuan, et al.
Publicado: (2026)
por: Ji, Kaixuan, et al.
Publicado: (2026)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
por: Ji, Kaixuan, et al.
Publicado: (2026)
por: Ji, Kaixuan, et al.
Publicado: (2026)
Feel-Good Thompson Sampling for Contextual Dueling Bandits
por: Li, Xuheng, et al.
Publicado: (2024)
por: Li, Xuheng, et al.
Publicado: (2024)
Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
por: Di, Qiwei, et al.
Publicado: (2024)
por: Di, Qiwei, et al.
Publicado: (2024)
Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence
por: Zhang, Shiyuan, et al.
Publicado: (2026)
por: Zhang, Shiyuan, et al.
Publicado: (2026)
Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
por: Di, Qiwei, et al.
Publicado: (2023)
por: Di, Qiwei, et al.
Publicado: (2023)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
por: Zhao, Qingyue, et al.
Publicado: (2026)
por: Zhao, Qingyue, et al.
Publicado: (2026)
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
por: Di, Qiwei, et al.
Publicado: (2024)
por: Di, Qiwei, et al.
Publicado: (2024)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
por: Zhao, Qingyue, et al.
Publicado: (2025)
por: Zhao, Qingyue, et al.
Publicado: (2025)
Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits
por: Li, Xuheng, et al.
Publicado: (2025)
por: Li, Xuheng, et al.
Publicado: (2025)
Understanding SGD with Exponential Moving Average: A Case Study in Linear Regression
por: Li, Xuheng, et al.
Publicado: (2025)
por: Li, Xuheng, et al.
Publicado: (2025)
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
por: Zhao, Heyang, et al.
Publicado: (2023)
por: Zhao, Heyang, et al.
Publicado: (2023)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
por: Di, Qiwei, et al.
Publicado: (2023)
por: Di, Qiwei, et al.
Publicado: (2023)
On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
por: Yu, Yue, et al.
Publicado: (2025)
por: Yu, Yue, et al.
Publicado: (2025)
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers
por: Li, Runjia, et al.
Publicado: (2024)
por: Li, Runjia, et al.
Publicado: (2024)
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
por: Zhao, Heyang, et al.
Publicado: (2024)
por: Zhao, Heyang, et al.
Publicado: (2024)
Reinforcement Learning from Human Feedback with Active Queries
por: Ji, Kaixuan, et al.
Publicado: (2024)
por: Ji, Kaixuan, et al.
Publicado: (2024)
Logarithmic Regret for Online KL-Regularized Reinforcement Learning
por: Zhao, Heyang, et al.
Publicado: (2025)
por: Zhao, Heyang, et al.
Publicado: (2025)
Minimax and Bayes Optimal Best-Arm Identification
por: Kato, Masahiro
Publicado: (2025)
por: Kato, Masahiro
Publicado: (2025)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
por: Yuan, Huizhuo, et al.
Publicado: (2024)
por: Yuan, Huizhuo, et al.
Publicado: (2024)
Relative Translation Invariant Wasserstein Distance
por: Wang, Binshuai, et al.
Publicado: (2024)
por: Wang, Binshuai, et al.
Publicado: (2024)
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
por: Chen, Zixiang, et al.
Publicado: (2024)
por: Chen, Zixiang, et al.
Publicado: (2024)
Minimax Limits of k-Fold Cross-Validation via Majority
por: Nachum, Ido, et al.
Publicado: (2026)
por: Nachum, Ido, et al.
Publicado: (2026)
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
por: Zhao, Heyang, et al.
Publicado: (2025)
por: Zhao, Heyang, et al.
Publicado: (2025)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
por: Huang, Audrey, et al.
Publicado: (2025)
por: Huang, Audrey, et al.
Publicado: (2025)
Minimax Optimal Simple Regret in Two-Armed Best-Arm Identification
por: Kato, Masahiro
Publicado: (2024)
por: Kato, Masahiro
Publicado: (2024)
Generalized Neyman Allocation for Locally Minimax Optimal Best-Arm Identification
por: Kato, Masahiro
Publicado: (2024)
por: Kato, Masahiro
Publicado: (2024)
Self-Play Preference Optimization for Language Model Alignment
por: Wu, Yue, et al.
Publicado: (2024)
por: Wu, Yue, et al.
Publicado: (2024)
Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
por: Lee, Harin, et al.
Publicado: (2026)
por: Lee, Harin, et al.
Publicado: (2026)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
por: Zhang, Junkai, et al.
Publicado: (2023)
por: Zhang, Junkai, et al.
Publicado: (2023)
Minimax Optimal Q Learning with Nearest Neighbors
por: Zhao, Puning, et al.
Publicado: (2023)
por: Zhao, Puning, et al.
Publicado: (2023)
Near Optimal Inference for the Best-Performing Algorithm
por: Painsky, Amichai
Publicado: (2025)
por: Painsky, Amichai
Publicado: (2025)
Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits
por: Hou, Yunlong, et al.
Publicado: (2024)
por: Hou, Yunlong, et al.
Publicado: (2024)
Minimax Optimality and Spectral Routing for Majority-Vote Ensembles under Markov Dependence
por: Shihab, Ibne Farabi, et al.
Publicado: (2026)
por: Shihab, Ibne Farabi, et al.
Publicado: (2026)
Best of Both Worlds: Regret Minimization versus Minimax Play
por: Müller, Adrian, et al.
Publicado: (2025)
por: Müller, Adrian, et al.
Publicado: (2025)
Minimax Optimal Algorithms with Fixed-$k$-Nearest Neighbors
por: Ryu, J. Jon, et al.
Publicado: (2022)
por: Ryu, J. Jon, et al.
Publicado: (2022)
MOSAIC: Minimax-Optimal Sparsity-Adaptive Inference for Change Points in Dynamic Networks
por: Fan, Yingying, et al.
Publicado: (2025)
por: Fan, Yingying, et al.
Publicado: (2025)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
por: Fan, Zhiyuan, et al.
Publicado: (2025)
por: Fan, Zhiyuan, et al.
Publicado: (2025)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
por: Kou, Yiwen, et al.
Publicado: (2024)
por: Kou, Yiwen, et al.
Publicado: (2024)
Minimax and Communication-Efficient Distributed Best Subset Selection with Oracle Property
por: Lan, Jingguo, et al.
Publicado: (2024)
por: Lan, Jingguo, et al.
Publicado: (2024)
Ejemplares similares
-
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
por: Ji, Kaixuan, et al.
Publicado: (2026) -
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
por: Ji, Kaixuan, et al.
Publicado: (2026) -
Feel-Good Thompson Sampling for Contextual Dueling Bandits
por: Li, Xuheng, et al.
Publicado: (2024) -
Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
por: Di, Qiwei, et al.
Publicado: (2024) -
Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence
por: Zhang, Shiyuan, et al.
Publicado: (2026)