Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
Fuente:
arXiv
Salvato in:
| Autori principali: | Wei, Wang, Yang, Tiankai, Chen, Hongjie, Zhao, Yue, Dernoncourt, Franck, Rossi, Ryan A., Eldardiry, Hoda |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Model Selection for Time Series Forecasting via LLMs
di: Wei, Wang, et al.
Pubblicazione: (2025)
di: Wei, Wang, et al.
Pubblicazione: (2025)
A Framework for Fine-Tuning LLMs using Heterogeneous Feedback
di: Aponte, Ryan, et al.
Pubblicazione: (2024)
di: Aponte, Ryan, et al.
Pubblicazione: (2024)
Sci-LoRA: Mixture of Scientific LoRAs for Cross-Domain Lay Paraphrasing
di: Cheng, Ming, et al.
Pubblicazione: (2025)
di: Cheng, Ming, et al.
Pubblicazione: (2025)
Multimodal Music Recommendation System using LLMs
di: Kandagatla, Srikar Prabhas, et al.
Pubblicazione: (2026)
di: Kandagatla, Srikar Prabhas, et al.
Pubblicazione: (2026)
Learning Emergence of Interaction Patterns across Independent RL Agents in Multi-Agent Environments
di: Baddam, Vasanth Reddy, et al.
Pubblicazione: (2024)
di: Baddam, Vasanth Reddy, et al.
Pubblicazione: (2024)
REFT: Resource-Efficient Federated Training Framework for Heterogeneous and Resource-Constrained Environments
di: Desai, Humaid Ahmed, et al.
Pubblicazione: (2023)
di: Desai, Humaid Ahmed, et al.
Pubblicazione: (2023)
Performative Prediction with Bandit Feedback: Learning through Reparameterization
di: Chen, Yatong, et al.
Pubblicazione: (2023)
di: Chen, Yatong, et al.
Pubblicazione: (2023)
Steering MoE LLMs via Expert (De)Activation
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2025)
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2025)
Forecasting Time Series with LLMs via Patch-Based Prompting and Decomposition
di: Bumb, Mayank, et al.
Pubblicazione: (2025)
di: Bumb, Mayank, et al.
Pubblicazione: (2025)
Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off
di: Li, Zhaochun, et al.
Pubblicazione: (2026)
di: Li, Zhaochun, et al.
Pubblicazione: (2026)
MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning
di: Tabassum, Afrina, et al.
Pubblicazione: (2025)
di: Tabassum, Afrina, et al.
Pubblicazione: (2025)
On Bits and Bandits: Quantifying the Regret-Information Trade-off
di: Shufaro, Itai, et al.
Pubblicazione: (2024)
di: Shufaro, Itai, et al.
Pubblicazione: (2024)
Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback
di: Li, Zitian, et al.
Pubblicazione: (2026)
di: Li, Zitian, et al.
Pubblicazione: (2026)
Clustering Items through Bandit Feedback: Finding the Right Feature out of Many
di: Graf, Maximilian, et al.
Pubblicazione: (2025)
di: Graf, Maximilian, et al.
Pubblicazione: (2025)
Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance
di: Gong, Ziling, et al.
Pubblicazione: (2026)
di: Gong, Ziling, et al.
Pubblicazione: (2026)
Lipschitz Bandits with Stochastic Delayed Feedback
di: Liu, Zhongxuan, et al.
Pubblicazione: (2025)
di: Liu, Zhongxuan, et al.
Pubblicazione: (2025)
GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs
di: Wang, Yue, et al.
Pubblicazione: (2025)
di: Wang, Yue, et al.
Pubblicazione: (2025)
Biased Dueling Bandits with Stochastic Delayed Feedback
di: Yi, Bongsoo, et al.
Pubblicazione: (2024)
di: Yi, Bongsoo, et al.
Pubblicazione: (2024)
Alleviating Community Fear in Disasters via Multi-Agent Actor-Critic Reinforcement Learning
di: Hakke, Yashodhan D., et al.
Pubblicazione: (2026)
di: Hakke, Yashodhan D., et al.
Pubblicazione: (2026)
Learning to Route and Schedule LLMs from User Retrials via Contextual Queueing Bandits
di: Bae, Seoungbin, et al.
Pubblicazione: (2026)
di: Bae, Seoungbin, et al.
Pubblicazione: (2026)
One Head, Many Models: Cross-Attention Routing for Cost-Aware LLM Selection
di: Pulishetty, Roshini, et al.
Pubblicazione: (2025)
di: Pulishetty, Roshini, et al.
Pubblicazione: (2025)
Best-of-Both-Worlds Policy Optimization for CMDPs with Bandit Feedback
di: Stradi, Francesco Emanuele, et al.
Pubblicazione: (2024)
di: Stradi, Francesco Emanuele, et al.
Pubblicazione: (2024)
Adaptive Client Sampling in Federated Learning via Online Learning with Bandit Feedback
di: Zhao, Boxin, et al.
Pubblicazione: (2021)
di: Zhao, Boxin, et al.
Pubblicazione: (2021)
Accuracy, Memory Efficiency and Generalization: A Comparative Study on Liquid Neural Networks and Recurrent Neural Networks
di: Zong, Shilong, et al.
Pubblicazione: (2025)
di: Zong, Shilong, et al.
Pubblicazione: (2025)
Speed Up the Cold-Start Learning in Two-Sided Bandits with Many Arms
di: Bayati, Mohsen, et al.
Pubblicazione: (2022)
di: Bayati, Mohsen, et al.
Pubblicazione: (2022)
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
di: Fernandez, Nigel, et al.
Pubblicazione: (2025)
di: Fernandez, Nigel, et al.
Pubblicazione: (2025)
Learning Equilibria in Matching Games with Bandit Feedback
di: Athanasopoulos, Andreas, et al.
Pubblicazione: (2025)
di: Athanasopoulos, Andreas, et al.
Pubblicazione: (2025)
Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts
di: Lotfi, Sanae, et al.
Pubblicazione: (2026)
di: Lotfi, Sanae, et al.
Pubblicazione: (2026)
Educating a Responsible AI Workforce: Piloting a Curricular Module on AI Policy in a Graduate Machine Learning Course
di: Weichert, James, et al.
Pubblicazione: (2025)
di: Weichert, James, et al.
Pubblicazione: (2025)
IPR: Intelligent Prompt Routing with User-Controlled Quality-Cost Trade-offs
di: Feng, Aosong, et al.
Pubblicazione: (2025)
di: Feng, Aosong, et al.
Pubblicazione: (2025)
Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion
di: Van Nguyen, Chien, et al.
Pubblicazione: (2026)
di: Van Nguyen, Chien, et al.
Pubblicazione: (2026)
Optimal Clustering with Bandit Feedback
di: Yang, Junwen, et al.
Pubblicazione: (2022)
di: Yang, Junwen, et al.
Pubblicazione: (2022)
Large Generative Graph Models
di: Wang, Yu, et al.
Pubblicazione: (2024)
di: Wang, Yu, et al.
Pubblicazione: (2024)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
di: Liu, Haolin, et al.
Pubblicazione: (2024)
di: Liu, Haolin, et al.
Pubblicazione: (2024)
Does Feedback Help in Bandits with Arm Erasures?
di: Karakas, Merve, et al.
Pubblicazione: (2025)
di: Karakas, Merve, et al.
Pubblicazione: (2025)
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2026)
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2026)
Nearest Neighbour with Bandit Feedback
di: Pasteris, Stephen, et al.
Pubblicazione: (2023)
di: Pasteris, Stephen, et al.
Pubblicazione: (2023)
Navigating Trade-offs: Policy Summarization for Multi-Objective Reinforcement Learning
di: Osika, Zuzanna, et al.
Pubblicazione: (2024)
di: Osika, Zuzanna, et al.
Pubblicazione: (2024)
Measuring Time-Series Dataset Similarity using Wasserstein Distance
di: Chen, Hongjie, et al.
Pubblicazione: (2025)
di: Chen, Hongjie, et al.
Pubblicazione: (2025)
Efficient Algorithms for Logistic Contextual Slate Bandits with Bandit Feedback
di: Goyal, Tanmay, et al.
Pubblicazione: (2025)
di: Goyal, Tanmay, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Efficient Model Selection for Time Series Forecasting via LLMs
di: Wei, Wang, et al.
Pubblicazione: (2025) -
A Framework for Fine-Tuning LLMs using Heterogeneous Feedback
di: Aponte, Ryan, et al.
Pubblicazione: (2024) -
Sci-LoRA: Mixture of Scientific LoRAs for Cross-Domain Lay Paraphrasing
di: Cheng, Ming, et al.
Pubblicazione: (2025) -
Multimodal Music Recommendation System using LLMs
di: Kandagatla, Srikar Prabhas, et al.
Pubblicazione: (2026) -
Learning Emergence of Interaction Patterns across Independent RL Agents in Multi-Agent Environments
di: Baddam, Vasanth Reddy, et al.
Pubblicazione: (2024)