ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Taberner-Miller, Annette |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Six Sigma Agent: Achieving Enterprise-Grade Reliability in LLM Systems Through Consensus-Driven Decomposed Execution
von: Patel, Khush, et al.
Veröffentlicht: (2026)
von: Patel, Khush, et al.
Veröffentlicht: (2026)
Federated Learning and Class Imbalances
von: Zhu, Siqi, et al.
Veröffentlicht: (2026)
von: Zhu, Siqi, et al.
Veröffentlicht: (2026)
Your Data, My Model: Learning Who Really Helps in Federated Learning
von: Abdurakhmanova, Shamsiiat, et al.
Veröffentlicht: (2024)
von: Abdurakhmanova, Shamsiiat, et al.
Veröffentlicht: (2024)
Dynamic Dual-Granularity Skill Bank for Agentic RL
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
von: Shekar, Pavan C, et al.
Veröffentlicht: (2025)
von: Shekar, Pavan C, et al.
Veröffentlicht: (2025)
Territory Paint Wars: Diagnosing and Mitigating Failure Modes in Competitive Multi-Agent PPO
von: Singh, Diyansha
Veröffentlicht: (2026)
von: Singh, Diyansha
Veröffentlicht: (2026)
RL-LLM-DT: An Automatic Decision Tree Generation Method Based on RL Evaluation and LLM Enhancement
von: Lin, Junjie, et al.
Veröffentlicht: (2024)
von: Lin, Junjie, et al.
Veröffentlicht: (2024)
Client-Conditional Federated Learning via Local Training Data Statistics
von: Brännvall, Rickard
Veröffentlicht: (2026)
von: Brännvall, Rickard
Veröffentlicht: (2026)
AI Agents: Evolution, Architecture, and Real-World Applications
von: Krishnan, Naveen
Veröffentlicht: (2025)
von: Krishnan, Naveen
Veröffentlicht: (2025)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
von: Jia, Xiao
Veröffentlicht: (2026)
von: Jia, Xiao
Veröffentlicht: (2026)
MACS: Multi-Agent Reinforcement Learning for Optimization of Crystal Structures
von: Zamaraeva, Elena, et al.
Veröffentlicht: (2025)
von: Zamaraeva, Elena, et al.
Veröffentlicht: (2025)
Dynamical Priors as a Training Objective in Reinforcement Learning
von: Subaharan, Sukesh
Veröffentlicht: (2026)
von: Subaharan, Sukesh
Veröffentlicht: (2026)
Semantic-Constrained Federated Aggregation: Convergence Theory and Privacy-Utility Bounds for Knowledge-Enhanced Distributed Learning
von: Arafat, Jahidul
Veröffentlicht: (2025)
von: Arafat, Jahidul
Veröffentlicht: (2025)
EARCP: Self-Regulating Coherence-Aware Ensemble Architecture for Sequential Decision Making -- Ensemble Auto-Regule par Coherence et Performance
von: Amega, Mike
Veröffentlicht: (2026)
von: Amega, Mike
Veröffentlicht: (2026)
Mechanical Conscience: A Mathematical Framework for Dependability of Machine Intelligenc
von: Batzorig, Munkhdegerekh, et al.
Veröffentlicht: (2026)
von: Batzorig, Munkhdegerekh, et al.
Veröffentlicht: (2026)
WorkflowGen:an adaptive workflow generation mechanism driven by trajectory experience
von: Wei, Ruocan, et al.
Veröffentlicht: (2026)
von: Wei, Ruocan, et al.
Veröffentlicht: (2026)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork
von: Jing, Yuheng, et al.
Veröffentlicht: (2026)
von: Jing, Yuheng, et al.
Veröffentlicht: (2026)
SPARK: Igniting Communication-Efficient Decentralized Learning via Stage-wise Projected NTK and Accelerated Regularization
von: Xia, Li
Veröffentlicht: (2025)
von: Xia, Li
Veröffentlicht: (2025)
Spatial-Temporal Learning-Based Distributed Routing for Dynamic LEO Satellite Networks
von: Chou, Po-Heng, et al.
Veröffentlicht: (2026)
von: Chou, Po-Heng, et al.
Veröffentlicht: (2026)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
The Efficiency Attenuation Phenomenon: A Computational Challenge to the Language of Thought Hypothesis
von: Zhang, Di
Veröffentlicht: (2026)
von: Zhang, Di
Veröffentlicht: (2026)
Generative Evolutionary Meta-Solver (GEMS): Scalable Surrogate-Free Multi-Agent Reinforcement Learning
von: Sharma, Alakh, et al.
Veröffentlicht: (2025)
von: Sharma, Alakh, et al.
Veröffentlicht: (2025)
SafetyDrift: Predicting When AI Agents Cross the Line Before They Actually Do
von: Dhodapkar, Aditya, et al.
Veröffentlicht: (2026)
von: Dhodapkar, Aditya, et al.
Veröffentlicht: (2026)
Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents
von: Nechepurenko, Maksym, et al.
Veröffentlicht: (2026)
von: Nechepurenko, Maksym, et al.
Veröffentlicht: (2026)
On the Fundamental Limitations of Decentralized Learnable Reward Shaping in Cooperative Multi-Agent Reinforcement Learning
von: Akella, Aditya
Veröffentlicht: (2025)
von: Akella, Aditya
Veröffentlicht: (2025)
Adaptive Latent-Space Constraints in Personalized Federated Learning
von: Ayromlou, Sana, et al.
Veröffentlicht: (2025)
von: Ayromlou, Sana, et al.
Veröffentlicht: (2025)
AdaptOrch: Task-Adaptive Multi-Agent Orchestration in the Era of LLM Performance Convergence
von: Yu, Geunbin
Veröffentlicht: (2026)
von: Yu, Geunbin
Veröffentlicht: (2026)
A General Framework for Off-Policy Learning with Partially-Observed Reward
von: Takehi, Rikiya, et al.
Veröffentlicht: (2025)
von: Takehi, Rikiya, et al.
Veröffentlicht: (2025)
Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training
von: Gu, Zhengyao, et al.
Veröffentlicht: (2026)
von: Gu, Zhengyao, et al.
Veröffentlicht: (2026)
Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents
von: Chen, Hao-Hsuan
Veröffentlicht: (2026)
von: Chen, Hao-Hsuan
Veröffentlicht: (2026)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
von: Pepe, Alberto, et al.
Veröffentlicht: (2026)
von: Pepe, Alberto, et al.
Veröffentlicht: (2026)
Bio AI Agent: A Multi-Agent Artificial Intelligence System for Autonomous CAR-T Cell Therapy Development with Integrated Target Discovery, Toxicity Prediction, and Rational Molecular Design
von: Ni, Yi, et al.
Veröffentlicht: (2025)
von: Ni, Yi, et al.
Veröffentlicht: (2025)
Mean-Field Reinforcement Learning without Synchrony
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory
von: Jiang, Rongjie, et al.
Veröffentlicht: (2026)
von: Jiang, Rongjie, et al.
Veröffentlicht: (2026)
Batched Nonparametric Bandits via k-Nearest Neighbor UCB
von: Arya, Sakshi
Veröffentlicht: (2025)
von: Arya, Sakshi
Veröffentlicht: (2025)
SCOPE: Selective Conformal Optimized Pairwise LLM Judging
von: Badshah, Sher, et al.
Veröffentlicht: (2026)
von: Badshah, Sher, et al.
Veröffentlicht: (2026)
Generating Realistic Safety-Critical Scenarios for Vehicle-Pedestrian Interactions
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
Sequential Monte Carlo Bandits
von: Urteaga, Iñigo, et al.
Veröffentlicht: (2018)
von: Urteaga, Iñigo, et al.
Veröffentlicht: (2018)
Multi-Agent Decision-Focused Learning via Value-Aware Sequential Communication
von: Amoh, Benjamin, et al.
Veröffentlicht: (2026)
von: Amoh, Benjamin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Six Sigma Agent: Achieving Enterprise-Grade Reliability in LLM Systems Through Consensus-Driven Decomposed Execution
von: Patel, Khush, et al.
Veröffentlicht: (2026) -
Federated Learning and Class Imbalances
von: Zhu, Siqi, et al.
Veröffentlicht: (2026) -
Your Data, My Model: Learning Who Really Helps in Federated Learning
von: Abdurakhmanova, Shamsiiat, et al.
Veröffentlicht: (2024) -
Dynamic Dual-Granularity Skill Bank for Agentic RL
von: Tu, Songjun, et al.
Veröffentlicht: (2026) -
Adaptive Minds: Empowering Agents with LoRA-as-Tools
von: Shekar, Pavan C, et al.
Veröffentlicht: (2025)