Flexible Empowerment at Reasoning with Extended Best-of-N Sampling
Fuente:
arXiv
Saved in:
| Main Author: | Kobayashi, Taisuke |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improvements of Dark Experience Replay and Reservoir Sampling towards Better Balance between Consolidation and Plasticity
by: Kobayashi, Taisuke
Published: (2025)
by: Kobayashi, Taisuke
Published: (2025)
Pseudo-Quantized Actor-Critic Algorithm for Robustness to Noisy Temporal Difference Error
by: Kobayashi, Taisuke
Published: (2026)
by: Kobayashi, Taisuke
Published: (2026)
Revisiting Experience Replayable Conditions
by: Kobayashi, Taisuke
Published: (2024)
by: Kobayashi, Taisuke
Published: (2024)
DROP: Distributional and Regular Optimism and Pessimism for Reinforcement Learning
by: Kobayashi, Taisuke
Published: (2024)
by: Kobayashi, Taisuke
Published: (2024)
Intentionally-underestimated Value Function at Terminal State for Temporal-difference Learning with Mis-designed Reward
by: Kobayashi, Taisuke
Published: (2023)
by: Kobayashi, Taisuke
Published: (2023)
Consolidated Adaptive T-soft Update for Deep Reinforcement Learning
by: Kobayashi, Taisuke
Published: (2022)
by: Kobayashi, Taisuke
Published: (2022)
CubeDAgger: Interactive Imitation Learning for Dynamic Systems with Efficient yet Low-risk Interaction
by: Kobayashi, Taisuke
Published: (2025)
by: Kobayashi, Taisuke
Published: (2025)
Variational Adaptive Noise and Dropout towards Stable Recurrent Neural Networks
by: Kobayashi, Taisuke, et al.
Published: (2025)
by: Kobayashi, Taisuke, et al.
Published: (2025)
Towards Autonomous Driving of Personal Mobility with Small and Noisy Dataset using Tsallis-statistics-based Behavioral Cloning
by: Kobayashi, Taisuke, et al.
Published: (2021)
by: Kobayashi, Taisuke, et al.
Published: (2021)
Design of Restricted Normalizing Flow towards Arbitrary Stochastic Policy with Computational Efficiency
by: Kobayashi, Taisuke, et al.
Published: (2024)
by: Kobayashi, Taisuke, et al.
Published: (2024)
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
by: Tang, Yung-Chen, et al.
Published: (2025)
by: Tang, Yung-Chen, et al.
Published: (2025)
The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
by: Bagirov, Farid, et al.
Published: (2025)
by: Bagirov, Farid, et al.
Published: (2025)
Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
by: Rho, Hyung Gyu, et al.
Published: (2025)
by: Rho, Hyung Gyu, et al.
Published: (2025)
Structured Pruning for Diverse Best-of-N Reasoning Optimization
by: Nguyen, Hieu Trung, et al.
Published: (2025)
by: Nguyen, Hieu Trung, et al.
Published: (2025)
Sharper Bounds for $\ell_p$ Sensitivity Sampling
by: Woodruff, David P., et al.
Published: (2023)
by: Woodruff, David P., et al.
Published: (2023)
Weber-Fechner Law in Temporal Difference learning derived from Control as Inference
by: Takahashi, Keiichiro, et al.
Published: (2024)
by: Takahashi, Keiichiro, et al.
Published: (2024)
Ridge Leverage Score Sampling for $\ell_p$ Subspace Approximation
by: Woodruff, David P., et al.
Published: (2024)
by: Woodruff, David P., et al.
Published: (2024)
BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
by: Gui, Lin, et al.
Published: (2024)
by: Gui, Lin, et al.
Published: (2024)
Stochastically Constrained Best Arm Identification with Thompson Sampling
by: Yang, Le, et al.
Published: (2025)
by: Yang, Le, et al.
Published: (2025)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
by: Chow, Yinlam, et al.
Published: (2024)
by: Chow, Yinlam, et al.
Published: (2024)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
by: Qiu, Jiahao, et al.
Published: (2024)
by: Qiu, Jiahao, et al.
Published: (2024)
Tight Sample Complexity Bounds for Entropic Best Policy Identification
by: Essakine, Amer, et al.
Published: (2026)
by: Essakine, Amer, et al.
Published: (2026)
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024)
by: Hughes, John, et al.
Published: (2024)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
by: Guo, Jizhou, et al.
Published: (2025)
by: Guo, Jizhou, et al.
Published: (2025)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
by: Huang, Audrey, et al.
Published: (2025)
by: Huang, Audrey, et al.
Published: (2025)
Majority of the Bests: Improving Best-of-N via Bootstrapping
by: Rakhsha, Amin, et al.
Published: (2025)
by: Rakhsha, Amin, et al.
Published: (2025)
Variational Best-of-N Alignment
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
by: Yu, Zhuohao, et al.
Published: (2026)
by: Yu, Zhuohao, et al.
Published: (2026)
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
by: Aminian, Gholamali, et al.
Published: (2025)
by: Aminian, Gholamali, et al.
Published: (2025)
Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
Enhancing Transfer Learning with Flexible Nonparametric Posterior Sampling
by: Lee, Hyungi, et al.
Published: (2024)
by: Lee, Hyungi, et al.
Published: (2024)
Universally Empowering Zeroth-Order Optimization via Adaptive Layer-wise Sampling
by: Wang, Fei, et al.
Published: (2026)
by: Wang, Fei, et al.
Published: (2026)
AdaBoN: Adaptive Best-of-N Alignment
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
Learning Generative Selection for Best-of-N
by: Toshniwal, Shubham, et al.
Published: (2026)
by: Toshniwal, Shubham, et al.
Published: (2026)
GenSelect: A Generative Approach to Best-of-N
by: Toshniwal, Shubham, et al.
Published: (2025)
by: Toshniwal, Shubham, et al.
Published: (2025)
BWS: Best Window Selection Based on Sample Scores for Data Pruning across Broad Ranges
by: Choi, Hoyong, et al.
Published: (2024)
by: Choi, Hoyong, et al.
Published: (2024)
Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
by: Bushipaka, Praveen, et al.
Published: (2025)
by: Bushipaka, Praveen, et al.
Published: (2025)
OkadaTorch: A Differentiable Programming of Okada Model to Calculate Displacements and Strains from Fault Parameters
by: Someya, Masayoshi, et al.
Published: (2025)
by: Someya, Masayoshi, et al.
Published: (2025)
Revisiting the (Sub)Optimality of Best-of-N for Inference-Time Alignment
by: Sriraman, Ved, et al.
Published: (2026)
by: Sriraman, Ved, et al.
Published: (2026)
Optimal Stopping vs Best-of-$N$ for Inference Time Optimization
by: Kalayci, Yusuf, et al.
Published: (2025)
by: Kalayci, Yusuf, et al.
Published: (2025)
Similar Items
-
Improvements of Dark Experience Replay and Reservoir Sampling towards Better Balance between Consolidation and Plasticity
by: Kobayashi, Taisuke
Published: (2025) -
Pseudo-Quantized Actor-Critic Algorithm for Robustness to Noisy Temporal Difference Error
by: Kobayashi, Taisuke
Published: (2026) -
Revisiting Experience Replayable Conditions
by: Kobayashi, Taisuke
Published: (2024) -
DROP: Distributional and Regular Optimism and Pessimism for Reinforcement Learning
by: Kobayashi, Taisuke
Published: (2024) -
Intentionally-underestimated Value Function at Terminal State for Temporal-difference Learning with Mis-designed Reward
by: Kobayashi, Taisuke
Published: (2023)