RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Geuter, Jonathan, Kornhardt, Gregor |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
por: Tang, Yung-Chen, et al.
Publicado: (2025)
por: Tang, Yung-Chen, et al.
Publicado: (2025)
AdaBoN: Adaptive Best-of-N Alignment
por: Raman, Vinod, et al.
Publicado: (2025)
por: Raman, Vinod, et al.
Publicado: (2025)
Universal Neural Optimal Transport
por: Geuter, Jonathan, et al.
Publicado: (2022)
por: Geuter, Jonathan, et al.
Publicado: (2022)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
por: Qiu, Jiahao, et al.
Publicado: (2024)
por: Qiu, Jiahao, et al.
Publicado: (2024)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
por: Huang, Audrey, et al.
Publicado: (2025)
por: Huang, Audrey, et al.
Publicado: (2025)
BoTTA: Benchmarking on-device Test Time Adaptation
por: Danilowski, Michal, et al.
Publicado: (2025)
por: Danilowski, Michal, et al.
Publicado: (2025)
BoSS: A Best-of-Strategies Selector as an Oracle for Deep Active Learning
por: Huseljic, Denis, et al.
Publicado: (2026)
por: Huseljic, Denis, et al.
Publicado: (2026)
Near-Optimal Online Deployment and Routing for Streaming LLMs
por: Li, Shaoang, et al.
Publicado: (2025)
por: Li, Shaoang, et al.
Publicado: (2025)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
por: Geuter, Jonathan, et al.
Publicado: (2025)
por: Geuter, Jonathan, et al.
Publicado: (2025)
A Regret Perspective on Online Multiple Testing
por: Hao, Qingyang, et al.
Publicado: (2026)
por: Hao, Qingyang, et al.
Publicado: (2026)
Best-of-$\infty$ -- Asymptotic Performance of Test-Time LLM Ensembling
por: Komiyama, Junpei, et al.
Publicado: (2025)
por: Komiyama, Junpei, et al.
Publicado: (2025)
BOND: Aligning LLMs with Best-of-N Distillation
por: Sessa, Pier Giuseppe, et al.
Publicado: (2024)
por: Sessa, Pier Giuseppe, et al.
Publicado: (2024)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
por: Lee, Jihoon, et al.
Publicado: (2025)
por: Lee, Jihoon, et al.
Publicado: (2025)
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
por: Zhang, Zeyu, et al.
Publicado: (2026)
por: Zhang, Zeyu, et al.
Publicado: (2026)
Revisiting the (Sub)Optimality of Best-of-N for Inference-Time Alignment
por: Sriraman, Ved, et al.
Publicado: (2026)
por: Sriraman, Ved, et al.
Publicado: (2026)
RoS-Guard: Robust and Scalable Online Change Detection with Delay-Optimal Guarantees
por: Zhu, Zelin, et al.
Publicado: (2025)
por: Zhu, Zelin, et al.
Publicado: (2025)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
por: Wang, Keyu, et al.
Publicado: (2025)
por: Wang, Keyu, et al.
Publicado: (2025)
S*: Test Time Scaling for Code Generation
por: Li, Dacheng, et al.
Publicado: (2025)
por: Li, Dacheng, et al.
Publicado: (2025)
T-POP: Test-Time Personalization with Online Preference Feedback
por: Qu, Zikun, et al.
Publicado: (2025)
por: Qu, Zikun, et al.
Publicado: (2025)
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
por: Xu, Zhiyuan, et al.
Publicado: (2026)
por: Xu, Zhiyuan, et al.
Publicado: (2026)
Constrained Best Arm Identification with Tests for Feasibility
por: Cai, Ting, et al.
Publicado: (2025)
por: Cai, Ting, et al.
Publicado: (2025)
Rethinking RoPE: A Mathematical Blueprint for N-dimensional Positional Embedding
por: Liu, Haiping, et al.
Publicado: (2025)
por: Liu, Haiping, et al.
Publicado: (2025)
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
por: Ding, Dujian, et al.
Publicado: (2025)
por: Ding, Dujian, et al.
Publicado: (2025)
Test-Time Training on Graphs with Large Language Models (LLMs)
por: Zhang, Jiaxin, et al.
Publicado: (2024)
por: Zhang, Jiaxin, et al.
Publicado: (2024)
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
por: Hübotter, Jonas, et al.
Publicado: (2024)
por: Hübotter, Jonas, et al.
Publicado: (2024)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
por: Miyamoto, Sora, et al.
Publicado: (2026)
por: Miyamoto, Sora, et al.
Publicado: (2026)
Understanding the Role of Training Data in Test-Time Scaling
por: Javanmard, Adel, et al.
Publicado: (2025)
por: Javanmard, Adel, et al.
Publicado: (2025)
Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs
por: Qiao, Ye, et al.
Publicado: (2025)
por: Qiao, Ye, et al.
Publicado: (2025)
Best-of-N Jailbreaking
por: Hughes, John, et al.
Publicado: (2024)
por: Hughes, John, et al.
Publicado: (2024)
Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
por: Qiao, Ye, et al.
Publicado: (2025)
por: Qiao, Ye, et al.
Publicado: (2025)
Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling
por: Tran, Dao, et al.
Publicado: (2026)
por: Tran, Dao, et al.
Publicado: (2026)
RoCA: Robust Contrastive One-class Time Series Anomaly Detection with Contaminated Data
por: Mou, Xudong, et al.
Publicado: (2025)
por: Mou, Xudong, et al.
Publicado: (2025)
Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
por: Rho, Hyung Gyu, et al.
Publicado: (2025)
por: Rho, Hyung Gyu, et al.
Publicado: (2025)
Variational Best-of-N Alignment
por: Amini, Afra, et al.
Publicado: (2024)
por: Amini, Afra, et al.
Publicado: (2024)
RouteLLM: Learning to Route LLMs with Preference Data
por: Ong, Isaac, et al.
Publicado: (2024)
por: Ong, Isaac, et al.
Publicado: (2024)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
por: Chen, Hao Mark, et al.
Publicado: (2025)
por: Chen, Hao Mark, et al.
Publicado: (2025)
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
por: Shen, Junhong, et al.
Publicado: (2025)
por: Shen, Junhong, et al.
Publicado: (2025)
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
por: Lee, Sanghyun, et al.
Publicado: (2025)
por: Lee, Sanghyun, et al.
Publicado: (2025)
$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models
por: Bilal, Ahsan, et al.
Publicado: (2026)
por: Bilal, Ahsan, et al.
Publicado: (2026)
Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
por: Kim, Yeongmin, et al.
Publicado: (2026)
por: Kim, Yeongmin, et al.
Publicado: (2026)
Ejemplares similares
-
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
por: Tang, Yung-Chen, et al.
Publicado: (2025) -
AdaBoN: Adaptive Best-of-N Alignment
por: Raman, Vinod, et al.
Publicado: (2025) -
Universal Neural Optimal Transport
por: Geuter, Jonathan, et al.
Publicado: (2022) -
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
por: Qiu, Jiahao, et al.
Publicado: (2024) -
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
por: Huang, Audrey, et al.
Publicado: (2025)