Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Biswas, Anish, Goel, Kanishk, S, Srivarshinee, Mohan, Jayashree, Khare, Alind, Parayil, Anjaly, Ramjee, Ramachandran, Bansal, Chetan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Niyama : Breaking the Silos of LLM Inference Serving
by: Goel, Kanishk, et al.
Published: (2025)
by: Goel, Kanishk, et al.
Published: (2025)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
by: Qiu, Haoran, et al.
Published: (2025)
by: Qiu, Haoran, et al.
Published: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring
by: Bastos, Anson, et al.
Published: (2026)
by: Bastos, Anson, et al.
Published: (2026)
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)
by: Parayil, Anjaly, et al.
Published: (2024)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
by: Kamath, Aditya K, et al.
Published: (2024)
by: Kamath, Aditya K, et al.
Published: (2024)
On Evaluating Performance of LLM Inference Serving Systems
by: Agrawal, Amey, et al.
Published: (2025)
by: Agrawal, Amey, et al.
Published: (2025)
Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
by: Gond, Raja, et al.
Published: (2025)
by: Gond, Raja, et al.
Published: (2025)
AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024)
by: Jain, Kunal, et al.
Published: (2024)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024)
by: Huang, Lexiang, et al.
Published: (2024)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
by: Deshmukh, Dhruv, et al.
Published: (2025)
by: Deshmukh, Dhruv, et al.
Published: (2025)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
by: Gond, Raja, et al.
Published: (2026)
by: Gond, Raja, et al.
Published: (2026)
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-Device Inference
by: Khare, Alind, et al.
Published: (2023)
by: Khare, Alind, et al.
Published: (2023)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
Intelligent Monitoring Framework for Cloud Services: A Data-Driven Approach
by: Srinivas, Pooja, et al.
Published: (2024)
by: Srinivas, Pooja, et al.
Published: (2024)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
BeLLMan: Controlling LLM Congestion
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
by: Reddy, Tella Rajashekhar, et al.
Published: (2026)
by: Reddy, Tella Rajashekhar, et al.
Published: (2026)
Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge
by: Koch, Fernando, et al.
Published: (2025)
by: Koch, Fernando, et al.
Published: (2025)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
by: Zhao, Xinkui, et al.
Published: (2025)
by: Zhao, Xinkui, et al.
Published: (2025)
Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems
by: Hussain, Fiza, et al.
Published: (2025)
by: Hussain, Fiza, et al.
Published: (2025)
Inference Load-Aware Orchestration for Hierarchical Federated Learning
by: Lackinger, Anna, et al.
Published: (2024)
by: Lackinger, Anna, et al.
Published: (2024)
Ensuring Fair LLM Serving Amid Diverse Applications
by: Khan, Redwan Ibne Seraj, et al.
Published: (2024)
by: Khan, Redwan Ibne Seraj, et al.
Published: (2024)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
by: Li, Maoliang, et al.
Published: (2026)
by: Li, Maoliang, et al.
Published: (2026)
Orchestrating Serverless Applications in the Edge Cloud Space Continuum: What Breaks and What is Next?
by: Malazi, Hadi Tabatabaee, et al.
Published: (2026)
by: Malazi, Hadi Tabatabaee, et al.
Published: (2026)
iDDS: Intelligent Distributed Dispatch and Scheduling for Workflow Orchestration
by: Guan, Wen, et al.
Published: (2025)
by: Guan, Wen, et al.
Published: (2025)
AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
by: Tokal, Shiva Sai Krishna Anand, et al.
Published: (2025)
by: Tokal, Shiva Sai Krishna Anand, et al.
Published: (2025)
MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
CUCo: An Agentic Framework for Compute and Communication Co-design
by: Hu, Bodun, et al.
Published: (2026)
by: Hu, Bodun, et al.
Published: (2026)
X-lifecycle Learning for Cloud Incident Management using LLMs
by: Goel, Drishti, et al.
Published: (2024)
by: Goel, Drishti, et al.
Published: (2024)
ASTRA: Accurate and Scalable ANNS-based Training of Extreme Classifiers
by: Mehta, Sonu, et al.
Published: (2024)
by: Mehta, Sonu, et al.
Published: (2024)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024)
by: Prabhu, Ramya, et al.
Published: (2024)
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Edge-Oriented Orchestration of Energy Services Using Graph-Driven Swarm Intelligence
by: Toderean, Liana, et al.
Published: (2026)
by: Toderean, Liana, et al.
Published: (2026)
LLM-assisted Agentic Edge Intelligence Framework
by: Dehury, Chinmaya Kumar, et al.
Published: (2026)
by: Dehury, Chinmaya Kumar, et al.
Published: (2026)
AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
by: Barke, Shraddha, et al.
Published: (2026)
by: Barke, Shraddha, et al.
Published: (2026)
Similar Items
-
Niyama : Breaking the Silos of LLM Inference Serving
by: Goel, Kanishk, et al.
Published: (2025) -
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
by: Qiu, Haoran, et al.
Published: (2025) -
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
by: Jaiswal, Shashwat, et al.
Published: (2025) -
A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring
by: Bastos, Anson, et al.
Published: (2026) -
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)