Reward-Based Online LLM Routing via NeuralUCB
Fuente:
arXiv
Guardado en:
| Autores principales: | Tsai, Ming-Hua, Tran, Phat |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards
por: Zhang, Dengjia, et al.
Publicado: (2026)
por: Zhang, Dengjia, et al.
Publicado: (2026)
Expected Reward Prediction, with Applications to Model Routing
por: Hasanaliyev, Kenan, et al.
Publicado: (2026)
por: Hasanaliyev, Kenan, et al.
Publicado: (2026)
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
por: Wang, Zige, et al.
Publicado: (2025)
por: Wang, Zige, et al.
Publicado: (2025)
Low-Resource Heuristics for Bahnaric Optical Character Recognition Improvement
por: Tran, Phat, et al.
Publicado: (2026)
por: Tran, Phat, et al.
Publicado: (2026)
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
por: Lu, Keming, et al.
Publicado: (2024)
por: Lu, Keming, et al.
Publicado: (2024)
Consolidating Rewarded Perturbations for LLM Post-Training
por: Zhang, Zheyu, et al.
Publicado: (2026)
por: Zhang, Zheyu, et al.
Publicado: (2026)
LLM Router: Rethinking Routing with Prefill Activations
por: Varshney, Tanay, et al.
Publicado: (2026)
por: Varshney, Tanay, et al.
Publicado: (2026)
Universal Model Routing for Efficient LLM Inference
por: Jitkrittum, Wittawat, et al.
Publicado: (2025)
por: Jitkrittum, Wittawat, et al.
Publicado: (2025)
Automated Rewards via LLM-Generated Progress Functions
por: Sarukkai, Vishnu, et al.
Publicado: (2024)
por: Sarukkai, Vishnu, et al.
Publicado: (2024)
Token-Level LLM Collaboration via FusionRoute
por: Xiong, Nuoya, et al.
Publicado: (2026)
por: Xiong, Nuoya, et al.
Publicado: (2026)
Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training
por: Xu, Ran, et al.
Publicado: (2026)
por: Xu, Ran, et al.
Publicado: (2026)
RouteLLM: Learning to Route LLMs with Preference Data
por: Ong, Isaac, et al.
Publicado: (2024)
por: Ong, Isaac, et al.
Publicado: (2024)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
por: Setlur, Amrith, et al.
Publicado: (2024)
por: Setlur, Amrith, et al.
Publicado: (2024)
UCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing
por: Kotte, Varun
Publicado: (2026)
por: Kotte, Varun
Publicado: (2026)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
por: Bai, Yang, et al.
Publicado: (2026)
por: Bai, Yang, et al.
Publicado: (2026)
Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards
por: Zhang, Yuxin, et al.
Publicado: (2025)
por: Zhang, Yuxin, et al.
Publicado: (2025)
TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
por: Zhang, Dan, et al.
Publicado: (2025)
por: Zhang, Dan, et al.
Publicado: (2025)
ReMoDetect: Reward Models Recognize Aligned LLM's Generations
por: Lee, Hyunseok, et al.
Publicado: (2024)
por: Lee, Hyunseok, et al.
Publicado: (2024)
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
por: Yang, Xinyi, et al.
Publicado: (2025)
por: Yang, Xinyi, et al.
Publicado: (2025)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
por: Behnam, Payman, et al.
Publicado: (2025)
por: Behnam, Payman, et al.
Publicado: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
por: Zhu, Xuekai, et al.
Publicado: (2025)
por: Zhu, Xuekai, et al.
Publicado: (2025)
TweakLLM: A Routing Architecture for Dynamic Tailoring of Cached Responses
por: Cheema, Muhammad Taha, et al.
Publicado: (2025)
por: Cheema, Muhammad Taha, et al.
Publicado: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
por: Dong, Hanze, et al.
Publicado: (2024)
por: Dong, Hanze, et al.
Publicado: (2024)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
por: Piękos, Piotr, et al.
Publicado: (2025)
por: Piękos, Piotr, et al.
Publicado: (2025)
Neural Bandit Based Optimal LLM Selection for a Pipeline of Subtasks
por: Atalar, Baran, et al.
Publicado: (2025)
por: Atalar, Baran, et al.
Publicado: (2025)
Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
por: Yu, Zichao, et al.
Publicado: (2025)
por: Yu, Zichao, et al.
Publicado: (2025)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
por: Nguyen, Hieu Trung, et al.
Publicado: (2026)
por: Nguyen, Hieu Trung, et al.
Publicado: (2026)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
por: Karlekar, Sweta, et al.
Publicado: (2026)
por: Karlekar, Sweta, et al.
Publicado: (2026)
Dr.LLM: Dynamic Layer Routing in LLMs
por: Heakl, Ahmed, et al.
Publicado: (2025)
por: Heakl, Ahmed, et al.
Publicado: (2025)
Energy-Based Reward Models for Robust Language Model Alignment
por: Lochab, Anamika, et al.
Publicado: (2025)
por: Lochab, Anamika, et al.
Publicado: (2025)
Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
por: Qiu, Wenjie, et al.
Publicado: (2025)
por: Qiu, Wenjie, et al.
Publicado: (2025)
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
por: Ding, Dujian, et al.
Publicado: (2025)
por: Ding, Dujian, et al.
Publicado: (2025)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
por: Xi, Zhiheng, et al.
Publicado: (2025)
por: Xi, Zhiheng, et al.
Publicado: (2025)
Reward-RAG: Enhancing RAG with Reward Driven Supervision
por: Nguyen, Thang, et al.
Publicado: (2024)
por: Nguyen, Thang, et al.
Publicado: (2024)
Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
por: Chuang, Yu-Neng, et al.
Publicado: (2025)
por: Chuang, Yu-Neng, et al.
Publicado: (2025)
RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
por: Jiang, Haoxiang, et al.
Publicado: (2026)
por: Jiang, Haoxiang, et al.
Publicado: (2026)
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
por: Wang, Haoxiang, et al.
Publicado: (2024)
por: Wang, Haoxiang, et al.
Publicado: (2024)
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
por: Ma, Qiyao, et al.
Publicado: (2026)
por: Ma, Qiyao, et al.
Publicado: (2026)
SELF: Self-Extend the Context Length With Logistic Growth Function
por: Dang, Phat Thanh, et al.
Publicado: (2025)
por: Dang, Phat Thanh, et al.
Publicado: (2025)
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
por: Li, Ming, et al.
Publicado: (2025)
por: Li, Ming, et al.
Publicado: (2025)
Ejemplares similares
-
SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards
por: Zhang, Dengjia, et al.
Publicado: (2026) -
Expected Reward Prediction, with Applications to Model Routing
por: Hasanaliyev, Kenan, et al.
Publicado: (2026) -
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
por: Wang, Zige, et al.
Publicado: (2025) -
Low-Resource Heuristics for Bahnaric Optical Character Recognition Improvement
por: Tran, Phat, et al.
Publicado: (2026) -
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
por: Lu, Keming, et al.
Publicado: (2024)