Scalable Best-of-N Selection for Large Language Models via Self-Certainty
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kang, Zhewei, Zhao, Xuandong, Song, Dawn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
von: Huang, Jiameng, et al.
Veröffentlicht: (2025)
von: Huang, Jiameng, et al.
Veröffentlicht: (2025)
Learning to Reason without External Rewards
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
Learning Generative Selection for Best-of-N
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2026)
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2026)
How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
Agent Instructs Large Language Models to be General Zero-Shot Reasoners
von: Crispino, Nicholas, et al.
Veröffentlicht: (2023)
von: Crispino, Nicholas, et al.
Veröffentlicht: (2023)
Re-Tuning: Overcoming the Compositionality Limits of Large Language Models with Recursive Tuning
von: Pasewark, Eric, et al.
Veröffentlicht: (2024)
von: Pasewark, Eric, et al.
Veröffentlicht: (2024)
Majority of the Bests: Improving Best-of-N via Bootstrapping
von: Rakhsha, Amin, et al.
Veröffentlicht: (2025)
von: Rakhsha, Amin, et al.
Veröffentlicht: (2025)
dLLM: Simple Diffusion Language Modeling
von: Zhou, Zhanhui, et al.
Veröffentlicht: (2026)
von: Zhou, Zhanhui, et al.
Veröffentlicht: (2026)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
Aligning Large Language Models by On-Policy Self-Judgment
von: Lee, Sangkyu, et al.
Veröffentlicht: (2024)
von: Lee, Sangkyu, et al.
Veröffentlicht: (2024)
SUV: Scalable Large Language Model Copyright Compliance with Regularized Selective Unlearning
von: Xu, Tianyang, et al.
Veröffentlicht: (2025)
von: Xu, Tianyang, et al.
Veröffentlicht: (2025)
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models
von: Yang, Yu, et al.
Veröffentlicht: (2024)
von: Yang, Yu, et al.
Veröffentlicht: (2024)
Best-of-N Jailbreaking
von: Hughes, John, et al.
Veröffentlicht: (2024)
von: Hughes, John, et al.
Veröffentlicht: (2024)
LLM-Select: Feature Selection with Large Language Models
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
Query-Conditioned Test-Time Self-Training for Large Language Models
von: Song, Chaehee, et al.
Veröffentlicht: (2026)
von: Song, Chaehee, et al.
Veröffentlicht: (2026)
Variational Best-of-N Alignment
von: Amini, Afra, et al.
Veröffentlicht: (2024)
von: Amini, Afra, et al.
Veröffentlicht: (2024)
Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
ProgCo: Program Helps Self-Correction of Large Language Models
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2025)
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2025)
Scalable Token-Level Hallucination Detection in Large Language Models
von: Min, Rui, et al.
Veröffentlicht: (2026)
von: Min, Rui, et al.
Veröffentlicht: (2026)
Stepwise Self-Consistent Mathematical Reasoning with Large Language Models
von: Zhao, Zilong, et al.
Veröffentlicht: (2024)
von: Zhao, Zilong, et al.
Veröffentlicht: (2024)
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
von: Zheng, Jinquan, et al.
Veröffentlicht: (2026)
von: Zheng, Jinquan, et al.
Veröffentlicht: (2026)
Selecting Large Language Model to Fine-tune via Rectified Scaling Law
von: Lin, Haowei, et al.
Veröffentlicht: (2024)
von: Lin, Haowei, et al.
Veröffentlicht: (2024)
DavIR: Data Selection via Implicit Reward for Large Language Models
von: Zhou, Haotian, et al.
Veröffentlicht: (2023)
von: Zhou, Haotian, et al.
Veröffentlicht: (2023)
Scalable Bayesian Low-Rank Adaptation of Large Language Models via Stochastic Variational Subspace Inference
von: Samplawski, Colin, et al.
Veröffentlicht: (2025)
von: Samplawski, Colin, et al.
Veröffentlicht: (2025)
AdaBoN: Adaptive Best-of-N Alignment
von: Raman, Vinod, et al.
Veröffentlicht: (2025)
von: Raman, Vinod, et al.
Veröffentlicht: (2025)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
von: Ji, Xiang, et al.
Veröffentlicht: (2024)
von: Ji, Xiang, et al.
Veröffentlicht: (2024)
Simple and Scalable Strategies to Continually Pre-train Large Language Models
von: Ibrahim, Adam, et al.
Veröffentlicht: (2024)
von: Ibrahim, Adam, et al.
Veröffentlicht: (2024)
SelfIE: Self-Interpretation of Large Language Model Embeddings
von: Chen, Haozhe, et al.
Veröffentlicht: (2024)
von: Chen, Haozhe, et al.
Veröffentlicht: (2024)
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
von: Li, Yun, et al.
Veröffentlicht: (2023)
von: Li, Yun, et al.
Veröffentlicht: (2023)
Auto-Evolve: Enhancing Large Language Model's Performance via Self-Reasoning Framework
von: Aswani, Krishna, et al.
Veröffentlicht: (2024)
von: Aswani, Krishna, et al.
Veröffentlicht: (2024)
LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tools and Self-Explanations
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
BOND: Aligning LLMs with Best-of-N Distillation
von: Sessa, Pier Giuseppe, et al.
Veröffentlicht: (2024)
von: Sessa, Pier Giuseppe, et al.
Veröffentlicht: (2024)
Differentially Private Zeroth-Order Methods for Scalable Large Language Model Finetuning
von: Liu, Z, et al.
Veröffentlicht: (2024)
von: Liu, Z, et al.
Veröffentlicht: (2024)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)
Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
An Undetectable Watermark for Generative Image Models
von: Gunn, Sam, et al.
Veröffentlicht: (2024)
von: Gunn, Sam, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025) -
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
von: Huang, Jiameng, et al.
Veröffentlicht: (2025) -
Learning to Reason without External Rewards
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025) -
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
von: Chow, Yinlam, et al.
Veröffentlicht: (2024) -
Learning Generative Selection for Best-of-N
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2026)