CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tang, Yung-Chen, Chen, Pin-Yu, Cavallaro, Andrea |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AdaBoN: Adaptive Best-of-N Alignment
par: Raman, Vinod, et autres
Publié: (2025)
par: Raman, Vinod, et autres
Publié: (2025)
RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
par: Geuter, Jonathan, et autres
Publié: (2025)
par: Geuter, Jonathan, et autres
Publié: (2025)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
par: Qiu, Jiahao, et autres
Publié: (2024)
par: Qiu, Jiahao, et autres
Publié: (2024)
Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
par: Rho, Hyung Gyu, et autres
Publié: (2025)
par: Rho, Hyung Gyu, et autres
Publié: (2025)
Defining and Evaluating Physical Safety for Large Language Models
par: Tang, Yung-Chen, et autres
Publié: (2024)
par: Tang, Yung-Chen, et autres
Publié: (2024)
Majority of the Bests: Improving Best-of-N via Bootstrapping
par: Rakhsha, Amin, et autres
Publié: (2025)
par: Rakhsha, Amin, et autres
Publié: (2025)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
par: Guo, Jizhou, et autres
Publié: (2025)
par: Guo, Jizhou, et autres
Publié: (2025)
BoSS: A Best-of-Strategies Selector as an Oracle for Deep Active Learning
par: Huseljic, Denis, et autres
Publié: (2026)
par: Huseljic, Denis, et autres
Publié: (2026)
Best-of-N Jailbreaking
par: Hughes, John, et autres
Publié: (2024)
par: Hughes, John, et autres
Publié: (2024)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
par: Huang, Audrey, et autres
Publié: (2025)
par: Huang, Audrey, et autres
Publié: (2025)
Computational Safety for Generative AI: A Signal Processing Perspective
par: Chen, Pin-Yu
Publié: (2025)
par: Chen, Pin-Yu
Publié: (2025)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
par: Chow, Yinlam, et autres
Publié: (2024)
par: Chow, Yinlam, et autres
Publié: (2024)
Variational Best-of-N Alignment
par: Amini, Afra, et autres
Publié: (2024)
par: Amini, Afra, et autres
Publié: (2024)
BoTTA: Benchmarking on-device Test Time Adaptation
par: Danilowski, Michal, et autres
Publié: (2025)
par: Danilowski, Michal, et autres
Publié: (2025)
BOND: Aligning LLMs with Best-of-N Distillation
par: Sessa, Pier Giuseppe, et autres
Publié: (2024)
par: Sessa, Pier Giuseppe, et autres
Publié: (2024)
Revisiting the (Sub)Optimality of Best-of-N for Inference-Time Alignment
par: Sriraman, Ved, et autres
Publié: (2026)
par: Sriraman, Ved, et autres
Publié: (2026)
Learning Generative Selection for Best-of-N
par: Toshniwal, Shubham, et autres
Publié: (2026)
par: Toshniwal, Shubham, et autres
Publié: (2026)
Data-Driven Lipschitz Continuity: A Cost-Effective Approach to Improve Adversarial Robustness
par: Chen, Erh-Chung, et autres
Publié: (2024)
par: Chen, Erh-Chung, et autres
Publié: (2024)
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
par: Hsieh, Cheng-Yu, et autres
Publié: (2024)
par: Hsieh, Cheng-Yu, et autres
Publié: (2024)
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles
par: Pukdee, Rattana, et autres
Publié: (2026)
par: Pukdee, Rattana, et autres
Publié: (2026)
Continual Learning for Adaptable Car-Following in Dynamic Traffic Environments
par: Chen, Xianda, et autres
Publié: (2024)
par: Chen, Xianda, et autres
Publié: (2024)
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
par: Chen, Feng, et autres
Publié: (2025)
par: Chen, Feng, et autres
Publié: (2025)
Minimal neuron ablation triggers catastrophic collapse in the language core of Large Vision-Language Models
par: Lu, Cen, et autres
Publié: (2025)
par: Lu, Cen, et autres
Publié: (2025)
Robust Calibration For Improved Weather Prediction Under Distributional Shift
par: Gilda, Sankalp, et autres
Publié: (2024)
par: Gilda, Sankalp, et autres
Publié: (2024)
Constrained Best Arm Identification with Tests for Feasibility
par: Cai, Ting, et autres
Publié: (2025)
par: Cai, Ting, et autres
Publié: (2025)
A Generative Car-following Model Conditioned On Driving Styles
par: Zhang, Yifan, et autres
Publié: (2021)
par: Zhang, Yifan, et autres
Publié: (2021)
MetaFollower: Adaptable Personalized Autonomous Car Following
par: Chen, Xianda, et autres
Publié: (2024)
par: Chen, Xianda, et autres
Publié: (2024)
Validity-Calibrated Reasoning Distillation
par: Saadi, Khouloud, et autres
Publié: (2026)
par: Saadi, Khouloud, et autres
Publié: (2026)
An Attention-Based Algorithm for Gravity Adaptation Zone Calibration
par: Yu, Chen
Publié: (2024)
par: Yu, Chen
Publié: (2024)
Two-Fidelity Best-Action Identification for Stochastic Minimax Tree
par: Chen, Peter, et autres
Publié: (2026)
par: Chen, Peter, et autres
Publié: (2026)
Test-time Diverse Reasoning by Riemannian Activation Steering
par: Khanh, Ly Tran Ho, et autres
Publié: (2025)
par: Khanh, Ly Tran Ho, et autres
Publié: (2025)
SLOT: Sample-specific Language Model Optimization at Test-time
par: Hu, Yang, et autres
Publié: (2025)
par: Hu, Yang, et autres
Publié: (2025)
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
par: Huang, Baihe, et autres
Publié: (2025)
par: Huang, Baihe, et autres
Publié: (2025)
Best-of-$\infty$ -- Asymptotic Performance of Test-Time LLM Ensembling
par: Komiyama, Junpei, et autres
Publié: (2025)
par: Komiyama, Junpei, et autres
Publié: (2025)
GREAT Score: Global Robustness Evaluation of Adversarial Perturbation using Generative Models
par: Li, Zaitang, et autres
Publié: (2023)
par: Li, Zaitang, et autres
Publié: (2023)
Maximizing Confidence Alone Improves Reasoning
par: Prabhudesai, Mihir, et autres
Publié: (2025)
par: Prabhudesai, Mihir, et autres
Publié: (2025)
Neural Clamping: Joint Input Perturbation and Temperature Scaling for Neural Network Calibration
par: Tang, Yung-Chen, et autres
Publié: (2022)
par: Tang, Yung-Chen, et autres
Publié: (2022)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
par: Kang, Zhewei, et autres
Publié: (2025)
par: Kang, Zhewei, et autres
Publié: (2025)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
par: Landesberg, Eddie
Publié: (2026)
par: Landesberg, Eddie
Publié: (2026)
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
par: Chen, Jinhao, et autres
Publié: (2025)
par: Chen, Jinhao, et autres
Publié: (2025)
Documents similaires
-
AdaBoN: Adaptive Best-of-N Alignment
par: Raman, Vinod, et autres
Publié: (2025) -
RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
par: Geuter, Jonathan, et autres
Publié: (2025) -
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
par: Qiu, Jiahao, et autres
Publié: (2024) -
Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
par: Rho, Hyung Gyu, et autres
Publié: (2025) -
Defining and Evaluating Physical Safety for Large Language Models
par: Tang, Yung-Chen, et autres
Publié: (2024)