Learning Generative Selection for Best-of-N
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Toshniwal, Shubham, Ficek, Aleksander, Jain, Siddhartha, Du, Wei, Noroozi, Vahid, Mahdavi, Sadegh, Majumdar, Somshubra, Gitman, Igor |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
par: Ficek, Aleksander, et autres
Publié: (2025)
par: Ficek, Aleksander, et autres
Publié: (2025)
GenSelect: A Generative Approach to Best-of-N
par: Toshniwal, Shubham, et autres
Publié: (2025)
par: Toshniwal, Shubham, et autres
Publié: (2025)
Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
par: Samadi, Mehrzad, et autres
Publié: (2025)
par: Samadi, Mehrzad, et autres
Publié: (2025)
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
par: Toshniwal, Shubham, et autres
Publié: (2024)
par: Toshniwal, Shubham, et autres
Publié: (2024)
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
par: Toshniwal, Shubham, et autres
Publié: (2024)
par: Toshniwal, Shubham, et autres
Publié: (2024)
Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
par: Noroozi, Vahid, et autres
Publié: (2024)
par: Noroozi, Vahid, et autres
Publié: (2024)
Scaling Generative Verifiers For Natural Language Mathematical Proof Verification And Selection
par: Mahdavi, Sadegh, et autres
Publié: (2025)
par: Mahdavi, Sadegh, et autres
Publié: (2025)
AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
par: Moshkov, Ivan, et autres
Publié: (2025)
par: Moshkov, Ivan, et autres
Publié: (2025)
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
par: Ahmad, Wasi Uddin, et autres
Publié: (2025)
par: Ahmad, Wasi Uddin, et autres
Publié: (2025)
OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique
par: Ahmad, Wasi Uddin, et autres
Publié: (2025)
par: Ahmad, Wasi Uddin, et autres
Publié: (2025)
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
par: Majumdar, Somshubra, et autres
Publié: (2024)
par: Majumdar, Somshubra, et autres
Publié: (2024)
Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision
par: Du, Wei, et autres
Publié: (2025)
par: Du, Wei, et autres
Publié: (2025)
OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
par: Ahmad, Wasi Uddin, et autres
Publié: (2025)
par: Ahmad, Wasi Uddin, et autres
Publié: (2025)
GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning
par: Ficek, Aleksander, et autres
Publié: (2024)
par: Ficek, Aleksander, et autres
Publié: (2024)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
par: Manikantan, Kawshik, et autres
Publié: (2024)
par: Manikantan, Kawshik, et autres
Publié: (2024)
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
par: Manikantan, Kawshik, et autres
Publié: (2024)
par: Manikantan, Kawshik, et autres
Publié: (2024)
Code Pretraining Improves Entity Tracking Abilities of Language Models
par: Kim, Najoung, et autres
Publié: (2024)
par: Kim, Najoung, et autres
Publié: (2024)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
par: Mahdavi, Sadegh, et autres
Publié: (2025)
par: Mahdavi, Sadegh, et autres
Publié: (2025)
Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection
par: Ma, Mingyu Derek, et autres
Publié: (2025)
par: Ma, Mingyu Derek, et autres
Publié: (2025)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
par: Kang, Zhewei, et autres
Publié: (2025)
par: Kang, Zhewei, et autres
Publié: (2025)
Best-of-N Jailbreaking
par: Hughes, John, et autres
Publié: (2024)
par: Hughes, John, et autres
Publié: (2024)
Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition
par: Noroozi, Vahid, et autres
Publié: (2023)
par: Noroozi, Vahid, et autres
Publié: (2023)
Variational Best-of-N Alignment
par: Amini, Afra, et autres
Publié: (2024)
par: Amini, Afra, et autres
Publié: (2024)
Lightweight reranking for language model generations
par: Jain, Siddhartha, et autres
Publié: (2023)
par: Jain, Siddhartha, et autres
Publié: (2023)
Majority of the Bests: Improving Best-of-N via Bootstrapping
par: Rakhsha, Amin, et autres
Publié: (2025)
par: Rakhsha, Amin, et autres
Publié: (2025)
AdaBoN: Adaptive Best-of-N Alignment
par: Raman, Vinod, et autres
Publié: (2025)
par: Raman, Vinod, et autres
Publié: (2025)
NeMo-Inspector: A Visualization Tool for LLM Generation Analysis
par: Gitman, Daria, et autres
Publié: (2025)
par: Gitman, Daria, et autres
Publié: (2025)
BOND: Aligning LLMs with Best-of-N Distillation
par: Sessa, Pier Giuseppe, et autres
Publié: (2024)
par: Sessa, Pier Giuseppe, et autres
Publié: (2024)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
par: Thrampoulidis, Christos, et autres
Publié: (2025)
par: Thrampoulidis, Christos, et autres
Publié: (2025)
Regret-Free Reinforcement Learning for LTL Specifications
par: Majumdar, Rupak, et autres
Publié: (2024)
par: Majumdar, Rupak, et autres
Publié: (2024)
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
par: Hu, Zhengyu, et autres
Publié: (2026)
par: Hu, Zhengyu, et autres
Publié: (2026)
ModeX: Evaluator-Free Best-of-N Selection for Open-Ended Generation
par: Choi, Hyeong Kyu, et autres
Publié: (2026)
par: Choi, Hyeong Kyu, et autres
Publié: (2026)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
par: Hong, Fenglu, et autres
Publié: (2025)
par: Hong, Fenglu, et autres
Publié: (2025)
Comparative Analysis of Demonstration Selection Algorithms for LLM In-Context Learning
par: Shu, Dong, et autres
Publié: (2024)
par: Shu, Dong, et autres
Publié: (2024)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
par: Qiu, Jiahao, et autres
Publié: (2024)
par: Qiu, Jiahao, et autres
Publié: (2024)
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models
par: Yang, Yu, et autres
Publié: (2024)
par: Yang, Yu, et autres
Publié: (2024)
Nemotron-4 340B Technical Report
par: Nvidia, et autres
Publié: (2024)
par: Nvidia, et autres
Publié: (2024)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
par: Chow, Yinlam, et autres
Publié: (2024)
par: Chow, Yinlam, et autres
Publié: (2024)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
par: Landesberg, Eddie
Publié: (2026)
par: Landesberg, Eddie
Publié: (2026)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
par: Guo, Jizhou, et autres
Publié: (2025)
par: Guo, Jizhou, et autres
Publié: (2025)
Documents similaires
-
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
par: Ficek, Aleksander, et autres
Publié: (2025) -
GenSelect: A Generative Approach to Best-of-N
par: Toshniwal, Shubham, et autres
Publié: (2025) -
Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
par: Samadi, Mehrzad, et autres
Publié: (2025) -
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
par: Toshniwal, Shubham, et autres
Publié: (2024) -
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
par: Toshniwal, Shubham, et autres
Publié: (2024)