Best-of-$\infty$ -- Asymptotic Performance of Test-Time LLM Ensembling
Fuente:
arXiv
Salvato in:
| Autori principali: | Komiyama, Junpei, Oba, Daisuke, Oyamada, Masafumi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Suboptimal Performance of the Bayes Optimal Algorithm in Frequentist Best Arm Identification
di: Komiyama, Junpei
Pubblicazione: (2022)
di: Komiyama, Junpei
Pubblicazione: (2022)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
di: Miyamoto, Sora, et al.
Pubblicazione: (2026)
di: Miyamoto, Sora, et al.
Pubblicazione: (2026)
CITE: Anytime-Valid Statistical Inference in LLM Self-Consistency
di: Ota, Hirofumi, et al.
Pubblicazione: (2026)
di: Ota, Hirofumi, et al.
Pubblicazione: (2026)
Jellyfish: A Large Language Model for Data Preprocessing
di: Zhang, Haochen, et al.
Pubblicazione: (2023)
di: Zhang, Haochen, et al.
Pubblicazione: (2023)
Rate-optimal Design for Anytime Best Arm Identification
di: Komiyama, Junpei, et al.
Pubblicazione: (2025)
di: Komiyama, Junpei, et al.
Pubblicazione: (2025)
Fixed Confidence Best Arm Identification in the Bayesian Setting
di: Jang, Kyoungseok, et al.
Pubblicazione: (2024)
di: Jang, Kyoungseok, et al.
Pubblicazione: (2024)
Replicability is Asymptotically Free in Multi-armed Bandits
di: Komiyama, Junpei, et al.
Pubblicazione: (2024)
di: Komiyama, Junpei, et al.
Pubblicazione: (2024)
RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
di: Geuter, Jonathan, et al.
Pubblicazione: (2025)
di: Geuter, Jonathan, et al.
Pubblicazione: (2025)
Constrained Best Arm Identification with Tests for Feasibility
di: Cai, Ting, et al.
Pubblicazione: (2025)
di: Cai, Ting, et al.
Pubblicazione: (2025)
High-dimensional Contextual Bandit Problem without Sparsity
di: Komiyama, Junpei, et al.
Pubblicazione: (2023)
di: Komiyama, Junpei, et al.
Pubblicazione: (2023)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
di: Huang, Audrey, et al.
Pubblicazione: (2025)
di: Huang, Audrey, et al.
Pubblicazione: (2025)
Auto-Prompt Ensemble for LLM Judge
di: Li, Jiajie, et al.
Pubblicazione: (2025)
di: Li, Jiajie, et al.
Pubblicazione: (2025)
DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance
di: Cohen, Seffi, et al.
Pubblicazione: (2025)
di: Cohen, Seffi, et al.
Pubblicazione: (2025)
CoSMo: a Framework to Instantiate Conditioned Process Simulation Models
di: Oyamada, Rafael S., et al.
Pubblicazione: (2023)
di: Oyamada, Rafael S., et al.
Pubblicazione: (2023)
LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents
di: Yano, Taro, et al.
Pubblicazione: (2025)
di: Yano, Taro, et al.
Pubblicazione: (2025)
Decocted Experience Improves Test-Time Inference in LLM Agents
di: Shen, Maohao, et al.
Pubblicazione: (2026)
di: Shen, Maohao, et al.
Pubblicazione: (2026)
Asymptotically Optimal Linear Best Feasible Arm Identification with Fixed Budget
di: Bian, Jie, et al.
Pubblicazione: (2025)
di: Bian, Jie, et al.
Pubblicazione: (2025)
Data-dependent Bounds with $T$-Optimal Best-of-Both-Worlds Guarantees in Multi-Armed Bandits using Stability-Penalty Matching
di: Nguyen, Quan, et al.
Pubblicazione: (2025)
di: Nguyen, Quan, et al.
Pubblicazione: (2025)
Deep Ensembles Secretly Perform Empirical Bayes
di: Loaiza-Ganem, Gabriel, et al.
Pubblicazione: (2025)
di: Loaiza-Ganem, Gabriel, et al.
Pubblicazione: (2025)
Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
di: Bushipaka, Praveen, et al.
Pubblicazione: (2025)
di: Bushipaka, Praveen, et al.
Pubblicazione: (2025)
Efficient Ensemble Conditional Independence Test Framework for Causal Discovery
di: Guan, Zhengkang, et al.
Pubblicazione: (2025)
di: Guan, Zhengkang, et al.
Pubblicazione: (2025)
STEB: In Search of the Best Evaluation Approach for Synthetic Time Series
di: Stenger, Michael, et al.
Pubblicazione: (2025)
di: Stenger, Michael, et al.
Pubblicazione: (2025)
Revisiting the (Sub)Optimality of Best-of-N for Inference-Time Alignment
di: Sriraman, Ved, et al.
Pubblicazione: (2026)
di: Sriraman, Ved, et al.
Pubblicazione: (2026)
Best-of-Tails: Bridging Optimism and Pessimism in Inference-Time Alignment
di: Hsu, Hsiang, et al.
Pubblicazione: (2026)
di: Hsu, Hsiang, et al.
Pubblicazione: (2026)
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
di: Tang, Yung-Chen, et al.
Pubblicazione: (2025)
di: Tang, Yung-Chen, et al.
Pubblicazione: (2025)
Generalization Performance of Ensemble Clustering: From Theory to Algorithm
di: Zhang, Xu, et al.
Pubblicazione: (2025)
di: Zhang, Xu, et al.
Pubblicazione: (2025)
Self-Improving LLM Agents at Test-Time
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2025)
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2025)
Sensi: Learn One Thing at a Time -- Curriculum-Based Test-Time Learning for LLM Game Agents
di: Arjmandi, Mohsen
Pubblicazione: (2026)
di: Arjmandi, Mohsen
Pubblicazione: (2026)
Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
di: Ishibashi, Yoichi, et al.
Pubblicazione: (2025)
di: Ishibashi, Yoichi, et al.
Pubblicazione: (2025)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
di: Wang, Yutong, et al.
Pubblicazione: (2025)
di: Wang, Yutong, et al.
Pubblicazione: (2025)
More Test-Time Compute Can Hurt: Overestimation Bias in LLM Beam Search
di: Dalal, Gal, et al.
Pubblicazione: (2026)
di: Dalal, Gal, et al.
Pubblicazione: (2026)
Finite-Time Regret Analysis of Retry-Aware Bandits
di: Tong, Bingkui, et al.
Pubblicazione: (2026)
di: Tong, Bingkui, et al.
Pubblicazione: (2026)
Improved Distribution Estimation in $\ell_\infty$
di: Cohen, Doron, et al.
Pubblicazione: (2026)
di: Cohen, Doron, et al.
Pubblicazione: (2026)
Adaptive Ensembles of Fine-Tuned Transformers for LLM-Generated Text Detection
di: Lai, Zhixin, et al.
Pubblicazione: (2024)
di: Lai, Zhixin, et al.
Pubblicazione: (2024)
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
di: Sareen, Kusha, et al.
Pubblicazione: (2025)
di: Sareen, Kusha, et al.
Pubblicazione: (2025)
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
di: Liu, Jia, et al.
Pubblicazione: (2025)
di: Liu, Jia, et al.
Pubblicazione: (2025)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
di: Zhu, Xiao, et al.
Pubblicazione: (2026)
di: Zhu, Xiao, et al.
Pubblicazione: (2026)
Atom of Thoughts for Markov LLM Test-Time Scaling
di: Teng, Fengwei, et al.
Pubblicazione: (2025)
di: Teng, Fengwei, et al.
Pubblicazione: (2025)
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
di: Monjur, Ocean, et al.
Pubblicazione: (2026)
di: Monjur, Ocean, et al.
Pubblicazione: (2026)
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
di: Ye, Wen, et al.
Pubblicazione: (2025)
di: Ye, Wen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Suboptimal Performance of the Bayes Optimal Algorithm in Frequentist Best Arm Identification
di: Komiyama, Junpei
Pubblicazione: (2022) -
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
di: Miyamoto, Sora, et al.
Pubblicazione: (2026) -
CITE: Anytime-Valid Statistical Inference in LLM Self-Consistency
di: Ota, Hirofumi, et al.
Pubblicazione: (2026) -
Jellyfish: A Large Language Model for Data Preprocessing
di: Zhang, Haochen, et al.
Pubblicazione: (2023) -
Rate-optimal Design for Anytime Best Arm Identification
di: Komiyama, Junpei, et al.
Pubblicazione: (2025)