The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Pai, Ning, Zhenyu, Zhao, Jieru, Cui, Weihao, Xu, Mengwei, Guo, Liwei, Chen, Xusheng, Shan, Yizhou |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RelServe: Fast LLM Inference Serving on Relational Data
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
TranSQL+: Serving Large Language Models with SQL on Low-Resource Hardware
by: Sun, Wenbo, et al.
Published: (2025)
by: Sun, Wenbo, et al.
Published: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Trinity: Disaggregating Vector Search from Prefill-Decode Disaggregation in LLM Serving
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
by: Kim, Jungwoo, et al.
Published: (2025)
by: Kim, Jungwoo, et al.
Published: (2025)
EdgeServe: A Streaming System for Decentralized Model Serving
by: Shaowang, Ted, et al.
Published: (2023)
by: Shaowang, Ted, et al.
Published: (2023)
HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving
by: Hu, Zhengding, et al.
Published: (2025)
by: Hu, Zhengding, et al.
Published: (2025)
Efficient LLM Serving for Agentic Workflows: A Data Systems Perspective
by: Wadlom, Noppanat, et al.
Published: (2026)
by: Wadlom, Noppanat, et al.
Published: (2026)
Serving Deep Learning Model in Relational Databases
by: Zhou, Lixi, et al.
Published: (2023)
by: Zhou, Lixi, et al.
Published: (2023)
Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
by: Wu, Fangzhou, et al.
Published: (2025)
by: Wu, Fangzhou, et al.
Published: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
by: Li, Suyi, et al.
Published: (2024)
by: Li, Suyi, et al.
Published: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
by: Liu, Banruo, et al.
Published: (2025)
by: Liu, Banruo, et al.
Published: (2025)
StreamTGN: A GPU-Efficient Serving System for Streaming Temporal Graph Neural Networks
by: Zhang, Lingling, et al.
Published: (2026)
by: Zhang, Lingling, et al.
Published: (2026)
Large Language Model Enhanced Text-to-SQL Generation: A Survey
by: Zhu, Xiaohu, et al.
Published: (2024)
by: Zhu, Xiaohu, et al.
Published: (2024)
Hispanex: Serving Libraries Serving Hispanics.
by: Aveney, Brian, et al.
Published: (1985)
by: Aveney, Brian, et al.
Published: (1985)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
RAGPulse: An Open-Source RAG Workload Trace to Optimize RAG Serving Systems
by: Wang, Zhengchao, et al.
Published: (2025)
by: Wang, Zhengchao, et al.
Published: (2025)
A CAP-like Trilemma for Large Language Models: Correctness, Non-bias, and Utility under Semantic Underdetermination
by: Venugopal, Vinu Ellampallil
Published: (2026)
by: Venugopal, Vinu Ellampallil
Published: (2026)
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
by: Liu, Yifei, et al.
Published: (2025)
by: Liu, Yifei, et al.
Published: (2025)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
by: Hu, Junhao, et al.
Published: (2024)
by: Hu, Junhao, et al.
Published: (2024)
Conceptual Schema Inference for Tabular Datasets using Large Language Models
by: Wu, Zhenyu, et al.
Published: (2026)
by: Wu, Zhenyu, et al.
Published: (2026)
LIRA: A Learning-based Query-aware Partition Framework for Large-scale ANN Search
by: Zeng, Ximu, et al.
Published: (2025)
by: Zeng, Ximu, et al.
Published: (2025)
Relational Database Augmented Large Language Model
by: Qin, Zongyue, et al.
Published: (2024)
by: Qin, Zongyue, et al.
Published: (2024)
Is Long Context All You Need? Leveraging LLM's Extended Context for NL2SQL
by: Chung, Yeounoh, et al.
Published: (2025)
by: Chung, Yeounoh, et al.
Published: (2025)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
by: Wu, Bingyang, et al.
Published: (2024)
by: Wu, Bingyang, et al.
Published: (2024)
Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching
by: Wang, Tianshu, et al.
Published: (2024)
by: Wang, Tianshu, et al.
Published: (2024)
RapidStore: An Efficient Dynamic Graph Storage System for Concurrent Queries
by: Hao, Chiyu, et al.
Published: (2025)
by: Hao, Chiyu, et al.
Published: (2025)
Conceptual Schema Inference for Tabular Datasets using Large Language Models
by: Wu, Zhenyu, et al.
Published: (2025)
by: Wu, Zhenyu, et al.
Published: (2025)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
by: Xie, Zhiqiang, et al.
Published: (2025)
by: Xie, Zhiqiang, et al.
Published: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
A Survey of LLM Inference Systems
by: Pan, James, et al.
Published: (2025)
by: Pan, James, et al.
Published: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
by: li, Fei, et al.
Published: (2026)
by: li, Fei, et al.
Published: (2026)
Can Large Language Models Be Query Optimizer for Relational Databases?
by: Tan, Jie, et al.
Published: (2025)
by: Tan, Jie, et al.
Published: (2025)
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
by: Yan, Jianxin, et al.
Published: (2025)
by: Yan, Jianxin, et al.
Published: (2025)
Efficient Data-aware Distance Comparison Operations for High-Dimensional Approximate Nearest Neighbor Search
by: Deng, Liwei, et al.
Published: (2024)
by: Deng, Liwei, et al.
Published: (2024)
ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries
by: Huang, Keke, et al.
Published: (2025)
by: Huang, Keke, et al.
Published: (2025)
K2: On Optimizing Distributed Transactions in a Multi-region Data Store with TrueTime Clocks (Extended Version)
by: Song, Haoze, et al.
Published: (2025)
by: Song, Haoze, et al.
Published: (2025)
Can Language Models Enable In-Context Database?
by: Pan, Yu, et al.
Published: (2024)
by: Pan, Yu, et al.
Published: (2024)
Halo: Domain-Aware Query Optimization for Long-Context Question Answering
by: Chunduri, Pramod, et al.
Published: (2026)
by: Chunduri, Pramod, et al.
Published: (2026)
Similar Items
-
RelServe: Fast LLM Inference Serving on Relational Data
by: Zhang, Xin, et al.
Published: (2025) -
TranSQL+: Serving Large Language Models with SQL on Low-Resource Hardware
by: Sun, Wenbo, et al.
Published: (2025) -
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024) -
Trinity: Disaggregating Vector Search from Prefill-Decode Disaggregation in LLM Serving
by: Liu, Yi, et al.
Published: (2025) -
Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
by: Kim, Jungwoo, et al.
Published: (2025)