Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
Fuente:
arXiv
Salvato in:
| Autori principali: | Rho, Hyung Gyu, Lee, Sian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Majority of the Bests: Improving Best-of-N via Bootstrapping
di: Rakhsha, Amin, et al.
Pubblicazione: (2025)
di: Rakhsha, Amin, et al.
Pubblicazione: (2025)
Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
di: Rho, Hyung Gyu
Pubblicazione: (2025)
di: Rho, Hyung Gyu
Pubblicazione: (2025)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
di: Guo, Jizhou, et al.
Pubblicazione: (2025)
di: Guo, Jizhou, et al.
Pubblicazione: (2025)
Learning Causal Structure of Time Series using Best Order Score Search
di: Mansilla, Irene Gema Castillo, et al.
Pubblicazione: (2026)
di: Mansilla, Irene Gema Castillo, et al.
Pubblicazione: (2026)
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
di: Tang, Yung-Chen, et al.
Pubblicazione: (2025)
di: Tang, Yung-Chen, et al.
Pubblicazione: (2025)
Generalized Neyman Allocation for Locally Minimax Optimal Best-Arm Identification
di: Kato, Masahiro
Pubblicazione: (2024)
di: Kato, Masahiro
Pubblicazione: (2024)
Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
di: Bushipaka, Praveen, et al.
Pubblicazione: (2025)
di: Bushipaka, Praveen, et al.
Pubblicazione: (2025)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
di: Chow, Yinlam, et al.
Pubblicazione: (2024)
di: Chow, Yinlam, et al.
Pubblicazione: (2024)
Conditional Generative Models are Sufficient to Sample from Any Causal Effect Estimand
di: Rahman, Md Musfiqur, et al.
Pubblicazione: (2024)
di: Rahman, Md Musfiqur, et al.
Pubblicazione: (2024)
Prediction-Powered Inference with Imputed Covariates and Nonuniform Sampling
di: Kluger, Dan M., et al.
Pubblicazione: (2025)
di: Kluger, Dan M., et al.
Pubblicazione: (2025)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
di: Qiu, Jiahao, et al.
Pubblicazione: (2024)
di: Qiu, Jiahao, et al.
Pubblicazione: (2024)
A Critical Perspective on Finite Sample Conformal Prediction Theory in Medical Applications
di: Kladny, Klaus-Rudolf, et al.
Pubblicazione: (2025)
di: Kladny, Klaus-Rudolf, et al.
Pubblicazione: (2025)
Best-of-N Jailbreaking
di: Hughes, John, et al.
Pubblicazione: (2024)
di: Hughes, John, et al.
Pubblicazione: (2024)
Is Elo Rating Reliable? A Study Under Model Misspecification
di: Tang, Shange, et al.
Pubblicazione: (2025)
di: Tang, Shange, et al.
Pubblicazione: (2025)
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles
di: Pukdee, Rattana, et al.
Pubblicazione: (2026)
di: Pukdee, Rattana, et al.
Pubblicazione: (2026)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
di: Huang, Audrey, et al.
Pubblicazione: (2025)
di: Huang, Audrey, et al.
Pubblicazione: (2025)
Recommending Best Paper Awards for ML/AI Conferences via the Isotonic Mechanism
di: Wen, Garrett G., et al.
Pubblicazione: (2026)
di: Wen, Garrett G., et al.
Pubblicazione: (2026)
Augmenting Limited and Biased RCTs through Pseudo-Sample Matching-Based Observational Data Fusion Method
di: Han, Kairong, et al.
Pubblicazione: (2025)
di: Han, Kairong, et al.
Pubblicazione: (2025)
Variational Best-of-N Alignment
di: Amini, Afra, et al.
Pubblicazione: (2024)
di: Amini, Afra, et al.
Pubblicazione: (2024)
Partial Identification Approach to Counterfactual Fairness Assessment
di: Rho, Saeyoung, et al.
Pubblicazione: (2025)
di: Rho, Saeyoung, et al.
Pubblicazione: (2025)
The Role of Contextual Information in Best Arm Identification
di: Kato, Masahiro, et al.
Pubblicazione: (2021)
di: Kato, Masahiro, et al.
Pubblicazione: (2021)
RCT Rejection Sampling for Causal Estimation Evaluation
di: Keith, Katherine A., et al.
Pubblicazione: (2023)
di: Keith, Katherine A., et al.
Pubblicazione: (2023)
AdaBoN: Adaptive Best-of-N Alignment
di: Raman, Vinod, et al.
Pubblicazione: (2025)
di: Raman, Vinod, et al.
Pubblicazione: (2025)
Learning Generative Selection for Best-of-N
di: Toshniwal, Shubham, et al.
Pubblicazione: (2026)
di: Toshniwal, Shubham, et al.
Pubblicazione: (2026)
Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees
di: Park, Sangwoo, et al.
Pubblicazione: (2025)
di: Park, Sangwoo, et al.
Pubblicazione: (2025)
PQMass: Probabilistic Assessment of the Quality of Generative Models using Probability Mass Estimation
di: Lemos, Pablo, et al.
Pubblicazione: (2024)
di: Lemos, Pablo, et al.
Pubblicazione: (2024)
Sample-Efficient Policy Space Response Oracles with Joint Experience Best Response
di: Bighashdel, Ariyan, et al.
Pubblicazione: (2026)
di: Bighashdel, Ariyan, et al.
Pubblicazione: (2026)
New Statistical Framework for Extreme Error Probability in High-Stakes Domains for Reliable Machine Learning
di: Michelucci, Umberto, et al.
Pubblicazione: (2025)
di: Michelucci, Umberto, et al.
Pubblicazione: (2025)
Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation
di: Martinon, Grégoire, et al.
Pubblicazione: (2026)
di: Martinon, Grégoire, et al.
Pubblicazione: (2026)
Revisiting the (Sub)Optimality of Best-of-N for Inference-Time Alignment
di: Sriraman, Ved, et al.
Pubblicazione: (2026)
di: Sriraman, Ved, et al.
Pubblicazione: (2026)
Efficient Causal Graph Discovery Using Large Language Models
di: Jiralerspong, Thomas, et al.
Pubblicazione: (2024)
di: Jiralerspong, Thomas, et al.
Pubblicazione: (2024)
BOND: Aligning LLMs with Best-of-N Distillation
di: Sessa, Pier Giuseppe, et al.
Pubblicazione: (2024)
di: Sessa, Pier Giuseppe, et al.
Pubblicazione: (2024)
Selection of the Most Probable Best
di: Kim, Taeho, et al.
Pubblicazione: (2022)
di: Kim, Taeho, et al.
Pubblicazione: (2022)
Statistical Limits and Efficient Algorithms for Differentially Private Federated Learning
di: Auddy, Arnab, et al.
Pubblicazione: (2026)
di: Auddy, Arnab, et al.
Pubblicazione: (2026)
First-Order Efficiency for Probabilistic Value Estimation via A Statistical Viewpoint
di: Liu, Ziqi, et al.
Pubblicazione: (2026)
di: Liu, Ziqi, et al.
Pubblicazione: (2026)
ABC3: Active Bayesian Causal Inference with Cohn Criteria in Randomized Experiments
di: Cha, Taehun, et al.
Pubblicazione: (2024)
di: Cha, Taehun, et al.
Pubblicazione: (2024)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
di: Kang, Zhewei, et al.
Pubblicazione: (2025)
di: Kang, Zhewei, et al.
Pubblicazione: (2025)
Modeling and Discovering Direct Causes for Predictive Models
di: Chen, Yizuo, et al.
Pubblicazione: (2024)
di: Chen, Yizuo, et al.
Pubblicazione: (2024)
CafeMed: Causal Attention Fusion Enhanced Medication Recommendation
di: Ren, Kelin, et al.
Pubblicazione: (2025)
di: Ren, Kelin, et al.
Pubblicazione: (2025)
Integrating Large Language Models in Causal Discovery: A Statistical Causal Approach
di: Takayama, Masayuki, et al.
Pubblicazione: (2024)
di: Takayama, Masayuki, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Majority of the Bests: Improving Best-of-N via Bootstrapping
di: Rakhsha, Amin, et al.
Pubblicazione: (2025) -
Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
di: Rho, Hyung Gyu
Pubblicazione: (2025) -
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
di: Guo, Jizhou, et al.
Pubblicazione: (2025) -
Learning Causal Structure of Time Series using Best Order Score Search
di: Mansilla, Irene Gema Castillo, et al.
Pubblicazione: (2026) -
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
di: Tang, Yung-Chen, et al.
Pubblicazione: (2025)