Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Ichihara, Yuki, Jinnai, Yuu, Morimura, Tetsuro, Ariu, Kaito, Abe, Kenshi, Sakamoto, Mitsuki, Uchibe, Eiji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment
by: Jinnai, Yuu, et al.
Published: (2024)
by: Jinnai, Yuu, et al.
Published: (2024)
Theoretical Guarantees for Minimum Bayes Risk Decoding
by: Ichihara, Yuki, et al.
Published: (2025)
by: Ichihara, Yuki, et al.
Published: (2025)
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024)
by: Morimura, Tetsuro, et al.
Published: (2024)
Model-Based Minimum Bayes Risk Decoding for Text Generation
by: Jinnai, Yuu, et al.
Published: (2023)
by: Jinnai, Yuu, et al.
Published: (2023)
Consensus Group Relative Policy Optimization for Text Generation
by: Ichihara, Yuki, et al.
Published: (2026)
by: Ichihara, Yuki, et al.
Published: (2026)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
by: Ichihara, Yuki, et al.
Published: (2025)
by: Ichihara, Yuki, et al.
Published: (2025)
Hyperparameter-Free Approach for Faster Minimum Bayes Risk Decoding
by: Jinnai, Yuu, et al.
Published: (2024)
by: Jinnai, Yuu, et al.
Published: (2024)
Boosting Perturbed Gradient Ascent for Last-Iterate Convergence in Games
by: Abe, Kenshi, et al.
Published: (2024)
by: Abe, Kenshi, et al.
Published: (2024)
Adaptively Perturbed Mirror Descent for Learning in Games
by: Abe, Kenshi, et al.
Published: (2023)
by: Abe, Kenshi, et al.
Published: (2023)
On the True Distribution Approximation of Minimum Bayes-Risk Decoding
by: Ohashi, Atsumoto, et al.
Published: (2024)
by: Ohashi, Atsumoto, et al.
Published: (2024)
Generating Diverse and High-Quality Texts by Minimum Bayes Risk Decoding
by: Jinnai, Yuu, et al.
Published: (2024)
by: Jinnai, Yuu, et al.
Published: (2024)
Does Cross-Cultural Alignment Change the Commonsense Morality of Language Models?
by: Jinnai, Yuu
Published: (2024)
by: Jinnai, Yuu
Published: (2024)
Asymmetric Perturbation in Solving Bilinear Saddle-Point Optimization
by: Abe, Kenshi, et al.
Published: (2025)
by: Abe, Kenshi, et al.
Published: (2025)
On the Power of Perturbation under Sampling in Solving Extensive-Form Games
by: Masaka, Wataru, et al.
Published: (2025)
by: Masaka, Wataru, et al.
Published: (2025)
Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts
by: Jinnai, Yuu, et al.
Published: (2024)
by: Jinnai, Yuu, et al.
Published: (2024)
Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport
by: Jinnai, Yuu
Published: (2025)
by: Jinnai, Yuu
Published: (2025)
Do Large Language Models Know Folktales? A Case Study of Yokai in Japanese Folktales
by: Tsutsumi, Ayuto, et al.
Published: (2025)
by: Tsutsumi, Ayuto, et al.
Published: (2025)
Return-Aligned Decision Transformer
by: Tanaka, Tsunehiko, et al.
Published: (2024)
by: Tanaka, Tsunehiko, et al.
Published: (2024)
Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
by: Jinnai, Yuu
Published: (2025)
by: Jinnai, Yuu
Published: (2025)
Last Iterate Convergence in Monotone Mean Field Games
by: Isobe, Noboru, et al.
Published: (2024)
by: Isobe, Noboru, et al.
Published: (2024)
Learning from Delayed Feedback in Games via Extra Prediction
by: Fujimoto, Yuma, et al.
Published: (2025)
by: Fujimoto, Yuma, et al.
Published: (2025)
Linear Convergence in Games with Delayed Feedback via Extra Prediction
by: Fujimoto, Yuma, et al.
Published: (2026)
by: Fujimoto, Yuma, et al.
Published: (2026)
Memory Asymmetry Creates Heteroclinic Orbits to Nash Equilibrium in Learning in Zero-Sum Games
by: Fujimoto, Yuma, et al.
Published: (2023)
by: Fujimoto, Yuma, et al.
Published: (2023)
Nash Equilibrium and Learning Dynamics in Three-Player Matching $m$-Action Games
by: Fujimoto, Yuma, et al.
Published: (2024)
by: Fujimoto, Yuma, et al.
Published: (2024)
Synchronization in Learning in Periodic Zero-Sum Games Triggers Divergence from Nash Equilibrium
by: Fujimoto, Yuma, et al.
Published: (2024)
by: Fujimoto, Yuma, et al.
Published: (2024)
Time-Varyingness in Auction Breaks Revenue Equivalence
by: Fujimoto, Yuma, et al.
Published: (2024)
by: Fujimoto, Yuma, et al.
Published: (2024)
Global Behavior of Learning Dynamics in Zero-Sum Games with Memory Asymmetry
by: Fujimoto, Yuma, et al.
Published: (2024)
by: Fujimoto, Yuma, et al.
Published: (2024)
Policy Testing in Markov Decision Processes
by: Ariu, Kaito, et al.
Published: (2025)
by: Ariu, Kaito, et al.
Published: (2025)
Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes
by: Morimura, Tetsuro, et al.
Published: (2022)
by: Morimura, Tetsuro, et al.
Published: (2022)
Reinforcement Learning for Edit-Based Non-Autoregressive Neural Machine Translation
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
by: Gui, Lin, et al.
Published: (2024)
by: Gui, Lin, et al.
Published: (2024)
Value-Based Large Language Model Agent Simulation for Mutual Evaluation of Trust and Interpersonal Closeness
by: Sakamoto, Yuki, et al.
Published: (2025)
by: Sakamoto, Yuki, et al.
Published: (2025)
Mean-Variance Efficient Reinforcement Learning with Applications to Dynamic Financial Investment
by: Kato, Masahiro, et al.
Published: (2020)
by: Kato, Masahiro, et al.
Published: (2020)
The Role of Contextual Information in Best Arm Identification
by: Kato, Masahiro, et al.
Published: (2021)
by: Kato, Masahiro, et al.
Published: (2021)
Variational Best-of-N Alignment
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
by: Chow, Yinlam, et al.
Published: (2024)
by: Chow, Yinlam, et al.
Published: (2024)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
by: Qiu, Jiahao, et al.
Published: (2024)
by: Qiu, Jiahao, et al.
Published: (2024)
AdaBoN: Adaptive Best-of-N Alignment
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
High-Fidelity Pseudo-label Generation by Large Language Models for Training Robust Radiology Report Classifiers
by: Wong, Brian, et al.
Published: (2025)
by: Wong, Brian, et al.
Published: (2025)
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
by: Beetham, James, et al.
Published: (2024)
by: Beetham, James, et al.
Published: (2024)
Similar Items
-
Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment
by: Jinnai, Yuu, et al.
Published: (2024) -
Theoretical Guarantees for Minimum Bayes Risk Decoding
by: Ichihara, Yuki, et al.
Published: (2025) -
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024) -
Model-Based Minimum Bayes Risk Decoding for Text Generation
by: Jinnai, Yuu, et al.
Published: (2023) -
Consensus Group Relative Policy Optimization for Text Generation
by: Ichihara, Yuki, et al.
Published: (2026)