OMEGA: A Low-Latency GNN Serving System for Large Graphs
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Geon-Woo, Kim, Donghyun, Moon, Jeongyoon, Liu, Henry, Khan, Tarannum, Iyer, Anand, Kim, Daehyeok, Akella, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
by: Hu, Bodun, et al.
Published: (2024)
by: Hu, Bodun, et al.
Published: (2024)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
by: Kim, Joon Ha, et al.
Published: (2026)
by: Kim, Joon Ha, et al.
Published: (2026)
Large Language Models as Realistic Microservice Trace Generators
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
Vulcan: Instance-Optimal Systems Heuristics Through LLM-Driven Search
by: Dwivedula, Rohit, et al.
Published: (2025)
by: Dwivedula, Rohit, et al.
Published: (2025)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
by: Wang, Shaoyu, et al.
Published: (2025)
by: Wang, Shaoyu, et al.
Published: (2025)
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
by: Dai, Yinwei, et al.
Published: (2023)
by: Dai, Yinwei, et al.
Published: (2023)
Man-Made Heuristics Are Dead. Long Live Code Generators!
by: Dwivedula, Rohit, et al.
Published: (2025)
by: Dwivedula, Rohit, et al.
Published: (2025)
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
by: Dai, Yinwei, et al.
Published: (2025)
by: Dai, Yinwei, et al.
Published: (2025)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
by: Wang, Wenfeng, et al.
Published: (2026)
by: Wang, Wenfeng, et al.
Published: (2026)
Software-Defined Agentic Serving
by: Agarwal, Saurabh, et al.
Published: (2026)
by: Agarwal, Saurabh, et al.
Published: (2026)
MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
by: Kwon, Miryeong, et al.
Published: (2025)
by: Kwon, Miryeong, et al.
Published: (2025)
ReInc: Scaling Training of Dynamic Graph Neural Networks
by: Guan, Mingyu, et al.
Published: (2025)
by: Guan, Mingyu, et al.
Published: (2025)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
by: Yuan, Yitao, et al.
Published: (2025)
by: Yuan, Yitao, et al.
Published: (2025)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
by: Kumar, Satyam, et al.
Published: (2026)
by: Kumar, Satyam, et al.
Published: (2026)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
by: Wang, Jun, et al.
Published: (2025)
by: Wang, Jun, et al.
Published: (2025)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
by: Yoon, Dongha, et al.
Published: (2025)
by: Yoon, Dongha, et al.
Published: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
Action Deviation-Aware Inference for Low-Latency Wireless Robots
by: Park, Jeyoung, et al.
Published: (2025)
by: Park, Jeyoung, et al.
Published: (2025)
Ripple: Scalable Incremental GNN Inferencing on Large Streaming Graphs
by: Naman, Pranjal, et al.
Published: (2025)
by: Naman, Pranjal, et al.
Published: (2025)
Formal Specification for Fast ACS: Low-Latency File-Based Ordered Message Delivery at Scale
by: Gupta, Sushant Kumar, et al.
Published: (2025)
by: Gupta, Sushant Kumar, et al.
Published: (2025)
Nalar: An agent serving framework
by: Laju, Marco, et al.
Published: (2026)
by: Laju, Marco, et al.
Published: (2026)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
by: Wu, Siyu, et al.
Published: (2025)
by: Wu, Siyu, et al.
Published: (2025)
HE2C: A Holistic Approach for Allocating Latency-Sensitive AI Tasks across Edge-Cloud
by: Kim, Minseo, et al.
Published: (2024)
by: Kim, Minseo, et al.
Published: (2024)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
by: Kim, Kihyun, et al.
Published: (2025)
by: Kim, Kihyun, et al.
Published: (2025)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
by: Niam, Arefin, et al.
Published: (2025)
by: Niam, Arefin, et al.
Published: (2025)
Incremental GNN Embedding Computation on Streaming Graphs
by: Wang, Qiange, et al.
Published: (2026)
by: Wang, Qiange, et al.
Published: (2026)
SpotVista: Availability-Aware Recommendation System for Reliable and Cost-Efficient Multi-Node Spot Instances
by: Kim, Taeyoon, et al.
Published: (2026)
by: Kim, Taeyoon, et al.
Published: (2026)
Cascadia: An Efficient Cascade Serving System for Large Language Models
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
Low-Latency Federated Fine-Tuning for Large Language Models Over Wireless Networks
by: Pang, Zhiwen, et al.
Published: (2026)
by: Pang, Zhiwen, et al.
Published: (2026)
Patchwork: A Unified Framework for RAG Serving
by: Hu, Bodun, et al.
Published: (2025)
by: Hu, Bodun, et al.
Published: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Large-Scale Graph Building in Dynamic Environments: Low Latency and High Quality
by: de Almeida, Filipe Miguel Gonçalves, et al.
Published: (2025)
by: de Almeida, Filipe Miguel Gonçalves, et al.
Published: (2025)
Efficient Graph-Based Approximate Nearest Neighbor Search Achieving: Low Latency Without Throughput Loss
by: Luo, Jingjia, et al.
Published: (2025)
by: Luo, Jingjia, et al.
Published: (2025)
ACE-GNN: Adaptive GNN Co-Inference with System-Aware Scheduling in Dynamic Edge Environments
by: Zhou, Ao, et al.
Published: (2025)
by: Zhou, Ao, et al.
Published: (2025)
Overcoming Latency-bound Limitations of Distributed Graph Algorithms using the HPX Runtime System
by: Mohammadiporshokooh, Karame, et al.
Published: (2026)
by: Mohammadiporshokooh, Karame, et al.
Published: (2026)
Hierarchical Autoscaling for Large Language Model Serving with Chiron
by: Patke, Archit, et al.
Published: (2025)
by: Patke, Archit, et al.
Published: (2025)
DeepVM: Integrating Spot and On-Demand VMs for Cost-Efficient Deep Learning Clusters in the Cloud
by: Kim, Yoochan, et al.
Published: (2024)
by: Kim, Yoochan, et al.
Published: (2024)
SDT-GNN: Streaming-based Distributed Training Framework for Graph Neural Networks
by: Huang, Xin, et al.
Published: (2024)
by: Huang, Xin, et al.
Published: (2024)
Accelerating Distributed Deep Learning using Lossless Homomorphic Compression
by: Li, Haoyu, et al.
Published: (2024)
by: Li, Haoyu, et al.
Published: (2024)
Similar Items
-
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
by: Hu, Bodun, et al.
Published: (2024) -
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
by: Kim, Joon Ha, et al.
Published: (2026) -
Large Language Models as Realistic Microservice Trace Generators
by: Kim, Donghyun, et al.
Published: (2024) -
Vulcan: Instance-Optimal Systems Heuristics Through LLM-Driven Search
by: Dwivedula, Rohit, et al.
Published: (2025) -
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
by: Wang, Shaoyu, et al.
Published: (2025)