Majority of the Bests: Improving Best-of-N via Bootstrapping
Fuente:
arXiv
Saved in:
| Main Authors: | Rakhsha, Amin, Madan, Kanika, Zhang, Tianyu, Farahmand, Amir-massoud, Khasahmadi, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PID Accelerated Temporal Difference Algorithms
by: Bedaywi, Mark, et al.
Published: (2024)
by: Bedaywi, Mark, et al.
Published: (2024)
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
by: Ma, Avery, et al.
Published: (2025)
by: Ma, Avery, et al.
Published: (2025)
Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
by: Rho, Hyung Gyu, et al.
Published: (2025)
by: Rho, Hyung Gyu, et al.
Published: (2025)
ALCM: Autonomous LLM-Augmented Causal Discovery Framework
by: Khatibi, Elahe, et al.
Published: (2024)
by: Khatibi, Elahe, et al.
Published: (2024)
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024)
by: Hughes, John, et al.
Published: (2024)
Deflated Dynamics Value Iteration
by: Lee, Jongmin, et al.
Published: (2024)
by: Lee, Jongmin, et al.
Published: (2024)
Variational Best-of-N Alignment
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Learning Generative Selection for Best-of-N
by: Toshniwal, Shubham, et al.
Published: (2026)
by: Toshniwal, Shubham, et al.
Published: (2026)
AdaBoN: Adaptive Best-of-N Alignment
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
BOND: Aligning LLMs with Best-of-N Distillation
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
by: Kang, Zhewei, et al.
Published: (2025)
by: Kang, Zhewei, et al.
Published: (2025)
Recommending Best Paper Awards for ML/AI Conferences via the Isotonic Mechanism
by: Wen, Garrett G., et al.
Published: (2026)
by: Wen, Garrett G., et al.
Published: (2026)
Dissecting Deep RL with High Update Ratios: Combatting Value Divergence
by: Hussing, Marcel, et al.
Published: (2024)
by: Hussing, Marcel, et al.
Published: (2024)
Learning Causal Structure of Time Series using Best Order Score Search
by: Mansilla, Irene Gema Castillo, et al.
Published: (2026)
by: Mansilla, Irene Gema Castillo, et al.
Published: (2026)
The Best Instruction-Tuning Data are Those That Fit
by: Zhang, Dylan, et al.
Published: (2025)
by: Zhang, Dylan, et al.
Published: (2025)
Text Rationalization for Robust Causal Effect Estimation
by: Zhang, Lijinghua, et al.
Published: (2025)
by: Zhang, Lijinghua, et al.
Published: (2025)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
by: Chow, Yinlam, et al.
Published: (2024)
by: Chow, Yinlam, et al.
Published: (2024)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
by: Landesberg, Eddie
Published: (2026)
by: Landesberg, Eddie
Published: (2026)
Generalized Neyman Allocation for Locally Minimax Optimal Best-Arm Identification
by: Kato, Masahiro
Published: (2024)
by: Kato, Masahiro
Published: (2024)
$λ$-models: Effective Decision-Aware Reinforcement Learning with Latent Models
by: Voelcker, Claas A, et al.
Published: (2023)
by: Voelcker, Claas A, et al.
Published: (2023)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
by: Qiu, Jiahao, et al.
Published: (2024)
by: Qiu, Jiahao, et al.
Published: (2024)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
by: Guo, Jizhou, et al.
Published: (2025)
by: Guo, Jizhou, et al.
Published: (2025)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
by: Zhang, Fred, et al.
Published: (2023)
by: Zhang, Fred, et al.
Published: (2023)
A Causal Lens for Evaluating Faithfulness Metrics
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
The Leaderboard Illusion
by: Singh, Shivalika, et al.
Published: (2025)
by: Singh, Shivalika, et al.
Published: (2025)
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
(Mis)Fitting: A Survey of Scaling Laws
by: Li, Margaret, et al.
Published: (2025)
by: Li, Margaret, et al.
Published: (2025)
RCT Rejection Sampling for Causal Estimation Evaluation
by: Keith, Katherine A., et al.
Published: (2023)
by: Keith, Katherine A., et al.
Published: (2023)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
by: Chew, Robert, et al.
Published: (2026)
by: Chew, Robert, et al.
Published: (2026)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
by: Boyeau, Pierre, et al.
Published: (2024)
by: Boyeau, Pierre, et al.
Published: (2024)
Efficient Exploration for LLMs
by: Dwaracherla, Vikranth, et al.
Published: (2024)
by: Dwaracherla, Vikranth, et al.
Published: (2024)
Adaptive Uncertainty Quantification for Generative AI
by: Kim, Jungeum, et al.
Published: (2024)
by: Kim, Jungeum, et al.
Published: (2024)
Industrial-Grade Smart Troubleshooting through Causal Technical Language Processing: a Proof of Concept
by: Trilla, Alexandre, et al.
Published: (2024)
by: Trilla, Alexandre, et al.
Published: (2024)
Propagation and Pitfalls: Reasoning-based Assessment of Knowledge Editing through Counterfactual Tasks
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning
by: Cai, Hengrui, et al.
Published: (2023)
by: Cai, Hengrui, et al.
Published: (2023)
CLEAR: Can Language Models Really Understand Causal Graphs?
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
Efficient Prompt Optimization Through the Lens of Best Arm Identification
by: Shi, Chengshuai, et al.
Published: (2024)
by: Shi, Chengshuai, et al.
Published: (2024)
s-ID: Causal Effect Identification in a Sub-Population
by: Abouei, Amir Mohammad, et al.
Published: (2023)
by: Abouei, Amir Mohammad, et al.
Published: (2023)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
by: Hong, Fenglu, et al.
Published: (2025)
by: Hong, Fenglu, et al.
Published: (2025)
Similar Items
-
PID Accelerated Temporal Difference Algorithms
by: Bedaywi, Mark, et al.
Published: (2024) -
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
by: Ma, Avery, et al.
Published: (2025) -
Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
by: Rho, Hyung Gyu, et al.
Published: (2025) -
ALCM: Autonomous LLM-Augmented Causal Discovery Framework
by: Khatibi, Elahe, et al.
Published: (2024) -
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024)