Which LLM to Play? Convergence-Aware Online Model Selection with Time-Increasing Bandits
Fuente:
arXiv
Salvato in:
| Autori principali: | Xia, Yu, Kong, Fang, Yu, Tong, Guo, Liya, Rossi, Ryan A., Kim, Sungchul, Li, Shuai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hallucination Diversity-Aware Active Learning for Text Summarization
di: Xia, Yu, et al.
Pubblicazione: (2024)
di: Xia, Yu, et al.
Pubblicazione: (2024)
Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution
di: Poon, Manhin, et al.
Pubblicazione: (2025)
di: Poon, Manhin, et al.
Pubblicazione: (2025)
Towards Improving Long-Tail Entity Predictions in Temporal Knowledge Graphs through Global Similarity and Weighted Sampling
di: Mirtaheri, Mehrnoosh, et al.
Pubblicazione: (2025)
di: Mirtaheri, Mehrnoosh, et al.
Pubblicazione: (2025)
Improved Bandits in Many-to-one Matching Markets with Incentive Compatibility
di: Kong, Fang, et al.
Pubblicazione: (2024)
di: Kong, Fang, et al.
Pubblicazione: (2024)
Finite-Time Regret Analysis of Retry-Aware Bandits
di: Tong, Bingkui, et al.
Pubblicazione: (2026)
di: Tong, Bingkui, et al.
Pubblicazione: (2026)
A Multi-LLM Debiasing Framework
di: Owens, Deonna M., et al.
Pubblicazione: (2024)
di: Owens, Deonna M., et al.
Pubblicazione: (2024)
PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMs
di: Hu, Xiaoyan, et al.
Pubblicazione: (2024)
di: Hu, Xiaoyan, et al.
Pubblicazione: (2024)
MAGNET: Autonomous Expert Model Generation via Decentralized Autoresearch and BitNet Training
di: Kim, Yongwan, et al.
Pubblicazione: (2026)
di: Kim, Yongwan, et al.
Pubblicazione: (2026)
Federated Large Language Models: Current Progress and Future Directions
di: Yao, Yuhang, et al.
Pubblicazione: (2024)
di: Yao, Yuhang, et al.
Pubblicazione: (2024)
Self-Play Enhancement via Advantage-Weighted Refinement in Online Federated LLM Fine-Tuning with Real-Time Feedback
di: Lee, Seohyun, et al.
Pubblicazione: (2026)
di: Lee, Seohyun, et al.
Pubblicazione: (2026)
Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent
di: Wu, Junda, et al.
Pubblicazione: (2025)
di: Wu, Junda, et al.
Pubblicazione: (2025)
Online Clustering of Dueling Bandits
di: Wang, Zhiyong, et al.
Pubblicazione: (2025)
di: Wang, Zhiyong, et al.
Pubblicazione: (2025)
A Multi-Armed Bandit Approach to Online Selection and Evaluation of Generative Models
di: Hu, Xiaoyan, et al.
Pubblicazione: (2024)
di: Hu, Xiaoyan, et al.
Pubblicazione: (2024)
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
di: Han, Zean, et al.
Pubblicazione: (2026)
di: Han, Zean, et al.
Pubblicazione: (2026)
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
di: Gallegos, Isabel O., et al.
Pubblicazione: (2024)
di: Gallegos, Isabel O., et al.
Pubblicazione: (2024)
Bias and Fairness in Large Language Models: A Survey
di: Gallegos, Isabel O., et al.
Pubblicazione: (2023)
di: Gallegos, Isabel O., et al.
Pubblicazione: (2023)
Visual Prompting in Multimodal Large Language Models: A Survey
di: Wu, Junda, et al.
Pubblicazione: (2024)
di: Wu, Junda, et al.
Pubblicazione: (2024)
Multi-Play Combinatorial Semi-Bandit Problem
di: Nakamura, Shintaro, et al.
Pubblicazione: (2025)
di: Nakamura, Shintaro, et al.
Pubblicazione: (2025)
From Selection to Generation: A Survey of LLM-based Active Learning
di: Xia, Yu, et al.
Pubblicazione: (2025)
di: Xia, Yu, et al.
Pubblicazione: (2025)
Efficient Model Selection for Time Series Forecasting via LLMs
di: Wei, Wang, et al.
Pubblicazione: (2025)
di: Wei, Wang, et al.
Pubblicazione: (2025)
Hybrid Combinatorial Multi-armed Bandits with Probabilistically Triggered Arms
di: Zhou, Kongchang, et al.
Pubblicazione: (2025)
di: Zhou, Kongchang, et al.
Pubblicazione: (2025)
Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits
di: Pershin, Maksim, et al.
Pubblicazione: (2026)
di: Pershin, Maksim, et al.
Pubblicazione: (2026)
KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
di: Ran, Dezhi, et al.
Pubblicazione: (2025)
di: Ran, Dezhi, et al.
Pubblicazione: (2025)
Training Robust Graph Neural Networks by Modeling Noise Dependencies
di: In, Yeonjun, et al.
Pubblicazione: (2025)
di: In, Yeonjun, et al.
Pubblicazione: (2025)
Play Style Identification Using Low-Level Representations of Play Traces in MicroRTS
di: Xia, Ruizhe Yu, et al.
Pubblicazione: (2025)
di: Xia, Ruizhe Yu, et al.
Pubblicazione: (2025)
Adversarial Bandit over Bandits: Hierarchical Bandits for Online Configuration Management
di: Avin, Chen, et al.
Pubblicazione: (2025)
di: Avin, Chen, et al.
Pubblicazione: (2025)
A Provably Convergent Plug-and-Play Framework for Stochastic Bilevel Optimization
di: Chu, Tianshu, et al.
Pubblicazione: (2025)
di: Chu, Tianshu, et al.
Pubblicazione: (2025)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
di: Liu, Hongyi, et al.
Pubblicazione: (2025)
di: Liu, Hongyi, et al.
Pubblicazione: (2025)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
di: Tan, Zhewen, et al.
Pubblicazione: (2026)
di: Tan, Zhewen, et al.
Pubblicazione: (2026)
Offline-to-Online Reinforcement Learning with Classifier-Free Diffusion Generation
di: Huang, Xiao, et al.
Pubblicazione: (2025)
di: Huang, Xiao, et al.
Pubblicazione: (2025)
Provably Convergent Primal-Dual DPO for Constrained LLM Alignment
di: Du, Yihan, et al.
Pubblicazione: (2025)
di: Du, Yihan, et al.
Pubblicazione: (2025)
Online Conformal Model Selection for Nonstationary Time Series
di: Li, Shibo, et al.
Pubblicazione: (2025)
di: Li, Shibo, et al.
Pubblicazione: (2025)
Cost-Effective Online Multi-LLM Selection with Versatile Reward Models
di: Dai, Xiangxiang, et al.
Pubblicazione: (2024)
di: Dai, Xiangxiang, et al.
Pubblicazione: (2024)
Knowledge-Aware Query Expansion with Large Language Models for Textual and Relational Retrieval
di: Xia, Yu, et al.
Pubblicazione: (2024)
di: Xia, Yu, et al.
Pubblicazione: (2024)
Causal Discovery in Semi-Stationary Time Series
di: Gao, Shanyun, et al.
Pubblicazione: (2024)
di: Gao, Shanyun, et al.
Pubblicazione: (2024)
A Federated Online Restless Bandit Framework for Cooperative Resource Allocation
di: Tong, Jingwen, et al.
Pubblicazione: (2024)
di: Tong, Jingwen, et al.
Pubblicazione: (2024)
Causal Discovery-Driven Change Point Detection in Time Series
di: Gao, Shanyun, et al.
Pubblicazione: (2024)
di: Gao, Shanyun, et al.
Pubblicazione: (2024)
R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM
di: Nguyen, Son, et al.
Pubblicazione: (2026)
di: Nguyen, Son, et al.
Pubblicazione: (2026)
RIE-Greedy: Regularization-Induced Exploration for Contextual Bandits
di: Li, Tong, et al.
Pubblicazione: (2026)
di: Li, Tong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Hallucination Diversity-Aware Active Learning for Text Summarization
di: Xia, Yu, et al.
Pubblicazione: (2024) -
Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution
di: Poon, Manhin, et al.
Pubblicazione: (2025) -
Towards Improving Long-Tail Entity Predictions in Temporal Knowledge Graphs through Global Similarity and Weighted Sampling
di: Mirtaheri, Mehrnoosh, et al.
Pubblicazione: (2025) -
Improved Bandits in Many-to-one Matching Markets with Incentive Compatibility
di: Kong, Fang, et al.
Pubblicazione: (2024) -
Finite-Time Regret Analysis of Retry-Aware Bandits
di: Tong, Bingkui, et al.
Pubblicazione: (2026)