Guardado en:
| Autores principales: | Sriraman, Ved, Block, Adam |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.05739 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
por: Huang, Audrey, et al.
Publicado: (2025)
por: Huang, Audrey, et al.
Publicado: (2025)
Best-of-Tails: Bridging Optimism and Pessimism in Inference-Time Alignment
por: Hsu, Hsiang, et al.
Publicado: (2026)
por: Hsu, Hsiang, et al.
Publicado: (2026)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
por: Qiu, Jiahao, et al.
Publicado: (2024)
por: Qiu, Jiahao, et al.
Publicado: (2024)
Variational Best-of-N Alignment
por: Amini, Afra, et al.
Publicado: (2024)
por: Amini, Afra, et al.
Publicado: (2024)
EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
por: Block, Adam, et al.
Publicado: (2025)
por: Block, Adam, et al.
Publicado: (2025)
AdaBoN: Adaptive Best-of-N Alignment
por: Raman, Vinod, et al.
Publicado: (2025)
por: Raman, Vinod, et al.
Publicado: (2025)
Real-Time Device Reach Forecasting Using HLL and MinHash Data Sketches
por: Muniyappa, Chandrashekar, et al.
Publicado: (2025)
por: Muniyappa, Chandrashekar, et al.
Publicado: (2025)
Learnable Chernoff Baselines for Inference-Time Alignment
por: Madhow, Sunil, et al.
Publicado: (2026)
por: Madhow, Sunil, et al.
Publicado: (2026)
Harnesses for Inference-Time Alignment over Execution Trajectories
por: Wang, Boyuan, et al.
Publicado: (2026)
por: Wang, Boyuan, et al.
Publicado: (2026)
Dynamic Search for Inference-Time Alignment in Diffusion Models
por: Li, Xiner, et al.
Publicado: (2025)
por: Li, Xiner, et al.
Publicado: (2025)
DAC-LoRA: Dynamic Adversarial Curriculum for Efficient and Robust Few-Shot Adaptation
por: Umrajkar, Ved
Publicado: (2025)
por: Umrajkar, Ved
Publicado: (2025)
Inference-Time Alignment of Diffusion Models with Direct Noise Optimization
por: Tang, Zhiwei, et al.
Publicado: (2024)
por: Tang, Zhiwei, et al.
Publicado: (2024)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
por: Chow, Yinlam, et al.
Publicado: (2024)
por: Chow, Yinlam, et al.
Publicado: (2024)
Compute Aligned Training: Optimizing for Test Time Inference
por: Ousherovitch, Adam, et al.
Publicado: (2026)
por: Ousherovitch, Adam, et al.
Publicado: (2026)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
por: Wang, Haichuan, et al.
Publicado: (2026)
por: Wang, Haichuan, et al.
Publicado: (2026)
Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
por: Jin, Luozhijie, et al.
Publicado: (2025)
por: Jin, Luozhijie, et al.
Publicado: (2025)
Best-of-N Jailbreaking
por: Hughes, John, et al.
Publicado: (2024)
por: Hughes, John, et al.
Publicado: (2024)
MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking
por: Zhao, Yizhou, et al.
Publicado: (2025)
por: Zhao, Yizhou, et al.
Publicado: (2025)
Is Behavior Cloning All You Need? Understanding Horizon in Imitation Learning
por: Foster, Dylan J., et al.
Publicado: (2024)
por: Foster, Dylan J., et al.
Publicado: (2024)
RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
por: Geuter, Jonathan, et al.
Publicado: (2025)
por: Geuter, Jonathan, et al.
Publicado: (2025)
Revisiting the Superficial Alignment Hypothesis
por: Raghavendra, Mohit, et al.
Publicado: (2024)
por: Raghavendra, Mohit, et al.
Publicado: (2024)
Majority of the Bests: Improving Best-of-N via Bootstrapping
por: Rakhsha, Amin, et al.
Publicado: (2025)
por: Rakhsha, Amin, et al.
Publicado: (2025)
GaussMark: A Practical Approach for Structural Watermarking of Language Models
por: Block, Adam, et al.
Publicado: (2025)
por: Block, Adam, et al.
Publicado: (2025)
Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
por: Rho, Hyung Gyu, et al.
Publicado: (2025)
por: Rho, Hyung Gyu, et al.
Publicado: (2025)
Active Learning via Regression Beyond Realizability
por: Ganju, Atul, et al.
Publicado: (2025)
por: Ganju, Atul, et al.
Publicado: (2025)
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
por: Tang, Yung-Chen, et al.
Publicado: (2025)
por: Tang, Yung-Chen, et al.
Publicado: (2025)
Experience-Guided Adaptation of Inference-Time Reasoning Strategies
por: Stein, Adam, et al.
Publicado: (2025)
por: Stein, Adam, et al.
Publicado: (2025)
Learning Generative Selection for Best-of-N
por: Toshniwal, Shubham, et al.
Publicado: (2026)
por: Toshniwal, Shubham, et al.
Publicado: (2026)
Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
por: Kanai, Sekitoshi, et al.
Publicado: (2025)
por: Kanai, Sekitoshi, et al.
Publicado: (2025)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
por: Bobbili, Sarat Chandra, et al.
Publicado: (2025)
por: Bobbili, Sarat Chandra, et al.
Publicado: (2025)
Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?
por: Gu, Zhuojun, et al.
Publicado: (2025)
por: Gu, Zhuojun, et al.
Publicado: (2025)
Best-of-$\infty$ -- Asymptotic Performance of Test-Time LLM Ensembling
por: Komiyama, Junpei, et al.
Publicado: (2025)
por: Komiyama, Junpei, et al.
Publicado: (2025)
STEB: In Search of the Best Evaluation Approach for Synthetic Time Series
por: Stenger, Michael, et al.
Publicado: (2025)
por: Stenger, Michael, et al.
Publicado: (2025)
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
por: Wang, Youkang, et al.
Publicado: (2026)
por: Wang, Youkang, et al.
Publicado: (2026)
BOND: Aligning LLMs with Best-of-N Distillation
por: Sessa, Pier Giuseppe, et al.
Publicado: (2024)
por: Sessa, Pier Giuseppe, et al.
Publicado: (2024)
GeSubNet: Gene Interaction Inference for Disease Subtype Network Generation
por: Yang, Ziwei, et al.
Publicado: (2024)
por: Yang, Ziwei, et al.
Publicado: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
por: Krishna, Kundan, et al.
Publicado: (2025)
por: Krishna, Kundan, et al.
Publicado: (2025)
Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review
por: Uehara, Masatoshi, et al.
Publicado: (2025)
por: Uehara, Masatoshi, et al.
Publicado: (2025)
Optimal Multi-Objective Best Arm Identification with Fixed Confidence
por: Chen, Zhirui, et al.
Publicado: (2025)
por: Chen, Zhirui, et al.
Publicado: (2025)
Ejemplares similares
-
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
por: Huang, Audrey, et al.
Publicado: (2025) -
Best-of-Tails: Bridging Optimism and Pessimism in Inference-Time Alignment
por: Hsu, Hsiang, et al.
Publicado: (2026) -
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
por: Qiu, Jiahao, et al.
Publicado: (2024) -
Variational Best-of-N Alignment
por: Amini, Afra, et al.
Publicado: (2024) -
EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
por: Block, Adam, et al.
Publicado: (2025)