RelServe: Fast LLM Inference Serving on Relational Data
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xin, Gao, Shihong, Shen, Yanyan, Li, Haoyang, Chen, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
by: Gao, Shihong, et al.
Published: (2025)
by: Gao, Shihong, et al.
Published: (2025)
Efficient LLM Serving for Agentic Workflows: A Data Systems Perspective
by: Wadlom, Noppanat, et al.
Published: (2026)
by: Wadlom, Noppanat, et al.
Published: (2026)
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
by: Zeng, Pai, et al.
Published: (2024)
by: Zeng, Pai, et al.
Published: (2024)
Rel: A Programming Language for Relational Data
by: Aref, Molham, et al.
Published: (2025)
by: Aref, Molham, et al.
Published: (2025)
Serving Deep Learning Model in Relational Databases
by: Zhou, Lixi, et al.
Published: (2023)
by: Zhou, Lixi, et al.
Published: (2023)
Trinity: Disaggregating Vector Search from Prefill-Decode Disaggregation in LLM Serving
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
by: Kim, Jungwoo, et al.
Published: (2025)
by: Kim, Jungwoo, et al.
Published: (2025)
HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving
by: Hu, Zhengding, et al.
Published: (2025)
by: Hu, Zhengding, et al.
Published: (2025)
LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round Annotation
by: Teng, Fei, et al.
Published: (2025)
by: Teng, Fei, et al.
Published: (2025)
EdgeServe: A Streaming System for Decentralized Model Serving
by: Shaowang, Ted, et al.
Published: (2023)
by: Shaowang, Ted, et al.
Published: (2023)
RelGNN: Composite Message Passing for Relational Deep Learning
by: Chen, Tianlang, et al.
Published: (2025)
by: Chen, Tianlang, et al.
Published: (2025)
Rel-MOSS: Towards Imbalanced Relational Deep Learning on Relational Databases
by: Yin, Jun, et al.
Published: (2026)
by: Yin, Jun, et al.
Published: (2026)
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
by: Kothapalli, Vignesh, et al.
Published: (2026)
by: Kothapalli, Vignesh, et al.
Published: (2026)
RelBench: A Benchmark for Deep Learning on Relational Databases
by: Robinson, Joshua, et al.
Published: (2024)
by: Robinson, Joshua, et al.
Published: (2024)
StreamTGN: A GPU-Efficient Serving System for Streaming Temporal Graph Neural Networks
by: Zhang, Lingling, et al.
Published: (2026)
by: Zhang, Lingling, et al.
Published: (2026)
Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
by: Wu, Fangzhou, et al.
Published: (2025)
by: Wu, Fangzhou, et al.
Published: (2025)
RAGPulse: An Open-Source RAG Workload Trace to Optimize RAG Serving Systems
by: Wang, Zhengchao, et al.
Published: (2025)
by: Wang, Zhengchao, et al.
Published: (2025)
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
by: Liu, Banruo, et al.
Published: (2025)
by: Liu, Banruo, et al.
Published: (2025)
Hispanex: Serving Libraries Serving Hispanics.
by: Aveney, Brian, et al.
Published: (1985)
by: Aveney, Brian, et al.
Published: (1985)
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
by: Yan, Jianxin, et al.
Published: (2026)
by: Yan, Jianxin, et al.
Published: (2026)
TranSQL+: Serving Large Language Models with SQL on Low-Resource Hardware
by: Sun, Wenbo, et al.
Published: (2025)
by: Sun, Wenbo, et al.
Published: (2025)
DataClaw: An Autonomous Data Agent with Instant Messaging Integration
by: Li, Huahang, et al.
Published: (2026)
by: Li, Huahang, et al.
Published: (2026)
Rel-HNN: Split Parallel Hypergraph Neural Network for Learning on Relational Databases
by: Alam, Md. Tanvir, et al.
Published: (2025)
by: Alam, Md. Tanvir, et al.
Published: (2025)
GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization
by: Tang, Tianhao, et al.
Published: (2026)
by: Tang, Tianhao, et al.
Published: (2026)
SiriusBI: A Comprehensive LLM-Powered Solution for Data Analytics in Business Intelligence
by: Jiang, Jie, et al.
Published: (2024)
by: Jiang, Jie, et al.
Published: (2024)
How Good Are Multi-dimensional Learned Indices? An Experimental Survey
by: Liu, Qiyu, et al.
Published: (2024)
by: Liu, Qiyu, et al.
Published: (2024)
A Survey of LLM Inference Systems
by: Pan, James, et al.
Published: (2025)
by: Pan, James, et al.
Published: (2025)
TRACE: A Time-Relational Approximate Cubing Engine for Fast Data Insights
by: Sivakumar, Suharsh, et al.
Published: (2024)
by: Sivakumar, Suharsh, et al.
Published: (2024)
Fast State Restoration in LLM Serving with HCache
by: Gao, Shiwei, et al.
Published: (2024)
by: Gao, Shiwei, et al.
Published: (2024)
A Query Optimization Method Utilizing Large Language Models
by: Yao, Zhiming, et al.
Published: (2025)
by: Yao, Zhiming, et al.
Published: (2025)
DAgent: A Relational Database-Driven Data Analysis Report Generation Agent
by: Xu, Wenyi, et al.
Published: (2025)
by: Xu, Wenyi, et al.
Published: (2025)
GPU-Accelerated Algorithms for Graph Vector Search: Taxonomy, Empirical Study, and Research Directions
by: Liu, Yaowen, et al.
Published: (2026)
by: Liu, Yaowen, et al.
Published: (2026)
Relational Database Distillation: From Structured Tables to Condensed Graph Data
by: Gao, Xinyi, et al.
Published: (2025)
by: Gao, Xinyi, et al.
Published: (2025)
Optimizing LLM Queries in Relational Data Analytics Workloads
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
FGIM: a Fast Graph-based Indexes Merging Framework for Approximate Nearest Neighbor Search
by: Wu, Zekai, et al.
Published: (2026)
by: Wu, Zekai, et al.
Published: (2026)
DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation
by: Fan, Meihao, et al.
Published: (2026)
by: Fan, Meihao, et al.
Published: (2026)
Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query Optimization
by: Zhu, Junhao, et al.
Published: (2025)
by: Zhu, Junhao, et al.
Published: (2025)
SeqRFM: Fast RFM Analysis in Sequence Data
by: Zheng, Yanxin, et al.
Published: (2024)
by: Zheng, Yanxin, et al.
Published: (2024)
Blitzcrank: Fast Semantic Compression for In-memory Online Transaction Processing
by: Qiao, Yiming, et al.
Published: (2024)
by: Qiao, Yiming, et al.
Published: (2024)
HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQL
by: Peng, You, et al.
Published: (2025)
by: Peng, You, et al.
Published: (2025)
Similar Items
-
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
by: Gao, Shihong, et al.
Published: (2025) -
Efficient LLM Serving for Agentic Workflows: A Data Systems Perspective
by: Wadlom, Noppanat, et al.
Published: (2026) -
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
by: Zeng, Pai, et al.
Published: (2024) -
Rel: A Programming Language for Relational Data
by: Aref, Molham, et al.
Published: (2025) -
Serving Deep Learning Model in Relational Databases
by: Zhou, Lixi, et al.
Published: (2023)