Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xutong, Atalar, Baran, Dai, Xiangxiang, Zuo, Jinhang, Wang, Siwei, Lui, John C. S., Chen, Wei, Joe-Wong, Carlee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Continuous Semantic Caching for Low-Cost LLM Serving
by: Atalar, Baran, et al.
Published: (2026)
by: Atalar, Baran, et al.
Published: (2026)
Offline Learning for Combinatorial Multi-armed Bandits
by: Liu, Xutong, et al.
Published: (2025)
by: Liu, Xutong, et al.
Published: (2025)
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
by: Dai, Xiangxiang, et al.
Published: (2025)
by: Dai, Xiangxiang, et al.
Published: (2025)
Batch-Size Independent Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms or Independent Arms
by: Liu, Xutong, et al.
Published: (2022)
by: Liu, Xutong, et al.
Published: (2022)
Neural Combinatorial Clustered Bandits for Recommendation Systems
by: Atalar, Baran, et al.
Published: (2024)
by: Atalar, Baran, et al.
Published: (2024)
Neural Bandit Based Optimal LLM Selection for a Pipeline of Subtasks
by: Atalar, Baran, et al.
Published: (2025)
by: Atalar, Baran, et al.
Published: (2025)
Cost-Effective Online Multi-LLM Selection with Versatile Reward Models
by: Dai, Xiangxiang, et al.
Published: (2024)
by: Dai, Xiangxiang, et al.
Published: (2024)
Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution
by: Poon, Manhin, et al.
Published: (2025)
by: Poon, Manhin, et al.
Published: (2025)
Intelligent Communication Planning for Constrained Environmental IoT Sensing with Reinforcement Learning
by: Hu, Yi, et al.
Published: (2023)
by: Hu, Yi, et al.
Published: (2023)
Offline Clustering of Linear Bandits: The Power of Clusters under Limited Data
by: Liu, Jingyuan, et al.
Published: (2025)
by: Liu, Jingyuan, et al.
Published: (2025)
Offline Clustering of Preference Learning with Active-data Augmentation
by: Liu, Jingyuan, et al.
Published: (2025)
by: Liu, Jingyuan, et al.
Published: (2025)
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
by: Liu, Xutong, et al.
Published: (2023)
by: Liu, Xutong, et al.
Published: (2023)
Tin-Tin: Towards Tiny Learning on Tiny Devices with Integer-based Neural Network Training
by: Hu, Yi, et al.
Published: (2025)
by: Hu, Yi, et al.
Published: (2025)
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE Inference
by: Han, Ziyi, et al.
Published: (2025)
by: Han, Ziyi, et al.
Published: (2025)
Stochastic Bandits Robust to Adversarial Attacks
by: Wang, Xuchuang, et al.
Published: (2024)
by: Wang, Xuchuang, et al.
Published: (2024)
CoRAST: Towards Foundation Model-Powered Correlated Data Analysis in Resource-Constrained CPS and IoT
by: Hu, Yi, et al.
Published: (2024)
by: Hu, Yi, et al.
Published: (2024)
Memory-Based Advantage Shaping for LLM-Guided Reinforcement Learning
by: Nourzad, Narjes, et al.
Published: (2026)
by: Nourzad, Narjes, et al.
Published: (2026)
MIRA: Memory-Integrated Reinforcement Learning Agent with Limited LLM Guidance
by: Nourzad, Narjes, et al.
Published: (2026)
by: Nourzad, Narjes, et al.
Published: (2026)
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond
by: Liu, Xutong, et al.
Published: (2024)
by: Liu, Xutong, et al.
Published: (2024)
Cost-Ordered Feasibility for Multi-Armed Bandits with Cost Subsidy
by: Juneja, Ishank, et al.
Published: (2026)
by: Juneja, Ishank, et al.
Published: (2026)
Combinatorial Logistic Bandits
by: Liu, Xutong, et al.
Published: (2024)
by: Liu, Xutong, et al.
Published: (2024)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Pairwise Elimination with Instance-Dependent Guarantees for Bandits with Cost Subsidy
by: Juneja, Ishank, et al.
Published: (2025)
by: Juneja, Ishank, et al.
Published: (2025)
HiLoRA: Adaptive Hierarchical LoRA Routing for Training-Free Domain Generalization
by: Han, Ziyi, et al.
Published: (2025)
by: Han, Ziyi, et al.
Published: (2025)
AxiomVision: Accuracy-Guaranteed Adaptive Visual Model Selection for Perspective-Aware Video Analytics
by: Dai, Xiangxiang, et al.
Published: (2024)
by: Dai, Xiangxiang, et al.
Published: (2024)
Federated Learning with Flexible Architectures
by: Park, Jong-Ik, et al.
Published: (2024)
by: Park, Jong-Ik, et al.
Published: (2024)
An LLM-Based Digital Twin for Optimizing Human-in-the Loop Systems
by: Yang, Hanqing, et al.
Published: (2024)
by: Yang, Hanqing, et al.
Published: (2024)
FedTLU: Federated Learning with Targeted Layer Updates
by: Park, Jong-Ik, et al.
Published: (2024)
by: Park, Jong-Ik, et al.
Published: (2024)
Demystifying Online Clustering of Bandits: Enhanced Exploration Under Stochastic and Smoothed Adversarial Contexts
by: Li, Zhuohua, et al.
Published: (2025)
by: Li, Zhuohua, et al.
Published: (2025)
M3Net: A Multi-Metric Mixture of Experts Network Digital Twin with Graph Neural Networks
by: Guda, Blessed, et al.
Published: (2025)
by: Guda, Blessed, et al.
Published: (2025)
DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration
by: Nourzad, Narjes, et al.
Published: (2025)
by: Nourzad, Narjes, et al.
Published: (2025)
Matrix-Space Reinforcement Learning for Reusing Local Transition Geometry
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
by: Qiao, Yifan, et al.
Published: (2024)
by: Qiao, Yifan, et al.
Published: (2024)
Quantum Algorithm for Online Exp-concave Optimization
by: He, Jianhao, et al.
Published: (2024)
by: He, Jianhao, et al.
Published: (2024)
A Multi-Agent Conversational Bandit Approach to Online Evaluation and Selection of User-Aligned LLM Responses
by: Dai, Xiangxiang, et al.
Published: (2025)
by: Dai, Xiangxiang, et al.
Published: (2025)
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving
by: Li, Shaoang, et al.
Published: (2026)
by: Li, Shaoang, et al.
Published: (2026)
FedSPD: A Soft-clustering Approach for Personalized Decentralized Federated Learning
by: Lin, I-Cheng, et al.
Published: (2024)
by: Lin, I-Cheng, et al.
Published: (2024)
GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching
by: Regmi, Sajal, et al.
Published: (2024)
by: Regmi, Sajal, et al.
Published: (2024)
Similar Items
-
Continuous Semantic Caching for Low-Cost LLM Serving
by: Atalar, Baran, et al.
Published: (2026) -
Offline Learning for Combinatorial Multi-armed Bandits
by: Liu, Xutong, et al.
Published: (2025) -
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
by: Dai, Xiangxiang, et al.
Published: (2025) -
Batch-Size Independent Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms or Independent Arms
by: Liu, Xutong, et al.
Published: (2022) -
Neural Combinatorial Clustered Bandits for Recommendation Systems
by: Atalar, Baran, et al.
Published: (2024)