Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Yinwei, Chen, Zhuofu, Iyer, Anand, Netravali, Ravi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
by: Dai, Yinwei, et al.
Published: (2023)
by: Dai, Yinwei, et al.
Published: (2023)
Kairos: A Scalable Serving System for Physical AI
by: Dai, Yinwei, et al.
Published: (2026)
by: Dai, Yinwei, et al.
Published: (2026)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
by: Pagonas, Nikos, et al.
Published: (2025)
by: Pagonas, Nikos, et al.
Published: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
by: Peng, You, et al.
Published: (2026)
by: Peng, You, et al.
Published: (2026)
Optimizing FaaS Platforms for MCP-enabled Agentic Workflows
by: Kulkarni, Varad, et al.
Published: (2026)
by: Kulkarni, Varad, et al.
Published: (2026)
AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
by: Tokal, Shiva Sai Krishna Anand, et al.
Published: (2025)
by: Tokal, Shiva Sai Krishna Anand, et al.
Published: (2025)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
by: Wagenländer, Marcel, et al.
Published: (2026)
by: Wagenländer, Marcel, et al.
Published: (2026)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
by: Du, Boxiao, et al.
Published: (2026)
by: Du, Boxiao, et al.
Published: (2026)
It Takes Two to Tango: Serverless Workflow Serving via Bilaterally Engaged Resource Adaptation
by: Wu, Jing, et al.
Published: (2025)
by: Wu, Jing, et al.
Published: (2025)
Workflow Mini-Apps: Portable, Scalable, Tunable & Faithful Representations of Scientific Workflows
by: Kilic, Ozgur Ozan, et al.
Published: (2024)
by: Kilic, Ozgur Ozan, et al.
Published: (2024)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
by: Zhang, Yuning, et al.
Published: (2026)
by: Zhang, Yuning, et al.
Published: (2026)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
Shelby: Decentralized Storage Designed to Serve
by: Goren, Guy, et al.
Published: (2025)
by: Goren, Guy, et al.
Published: (2025)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
by: Bu, Tianci, et al.
Published: (2026)
by: Bu, Tianci, et al.
Published: (2026)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
by: Yoshimura, Takeshi, et al.
Published: (2026)
by: Yoshimura, Takeshi, et al.
Published: (2026)
Wherefore Art Thou? Provenance-Guided Automatic Online Debugging with Lumos
by: Chen, Jingyuan, et al.
Published: (2026)
by: Chen, Jingyuan, et al.
Published: (2026)
Batch Query Processing and Optimization for Agentic Workflows
by: Shen, Junyi, et al.
Published: (2025)
by: Shen, Junyi, et al.
Published: (2025)
Software-Defined Agentic Serving
by: Agarwal, Saurabh, et al.
Published: (2026)
by: Agarwal, Saurabh, et al.
Published: (2026)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
by: Zhong, Yinmin, et al.
Published: (2024)
by: Zhong, Yinmin, et al.
Published: (2024)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
by: Xiang, Yuxing, et al.
Published: (2025)
by: Xiang, Yuxing, et al.
Published: (2025)
PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
by: Fang, Fei, et al.
Published: (2025)
by: Fang, Fei, et al.
Published: (2025)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
by: Kim, Geon-Woo, et al.
Published: (2025)
by: Kim, Geon-Woo, et al.
Published: (2025)
exa-AMD: A Scalable Workflow for Accelerating AI-Assisted Materials Discovery and Design
by: Moraru, Maxim, et al.
Published: (2025)
by: Moraru, Maxim, et al.
Published: (2025)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
by: Qiu, Haoran, et al.
Published: (2025)
by: Qiu, Haoran, et al.
Published: (2025)
Hierarchical Autoscaling for Large Language Model Serving with Chiron
by: Patke, Archit, et al.
Published: (2025)
by: Patke, Archit, et al.
Published: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
by: Ruan, Chaoyi, et al.
Published: (2025)
by: Ruan, Chaoyi, et al.
Published: (2025)
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
by: Heidari, Sina, et al.
Published: (2026)
by: Heidari, Sina, et al.
Published: (2026)
LegoDiffusion: Micro-Serving Text-to-Image Diffusion Workflows
by: Yang, Lingyun, et al.
Published: (2026)
by: Yang, Lingyun, et al.
Published: (2026)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
MoLink: Distributed and Efficient Serving Framework for Large Models
by: Jin, Lewei, et al.
Published: (2025)
by: Jin, Lewei, et al.
Published: (2025)
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
KS+: Predicting Workflow Task Memory Usage Over Time
by: Bader, Jonathan, et al.
Published: (2024)
by: Bader, Jonathan, et al.
Published: (2024)
Enabling Elastic Model Serving with MultiWorld
by: Lee, Myungjin, et al.
Published: (2024)
by: Lee, Myungjin, et al.
Published: (2024)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
by: Du, Jiangsu, et al.
Published: (2025)
by: Du, Jiangsu, et al.
Published: (2025)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
by: Bai, Fengyao, et al.
Published: (2026)
by: Bai, Fengyao, et al.
Published: (2026)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
by: Ye, Fanjiang, et al.
Published: (2026)
by: Ye, Fanjiang, et al.
Published: (2026)
Similar Items
-
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
by: Dai, Yinwei, et al.
Published: (2023) -
Kairos: A Scalable Serving System for Physical AI
by: Dai, Yinwei, et al.
Published: (2026) -
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
by: Pan, Rui, et al.
Published: (2025) -
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
by: Pagonas, Nikos, et al.
Published: (2025) -
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
by: Peng, You, et al.
Published: (2026)