HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, You, Jiang, Youhe, Li, Wenshuang, Xu, Xu, Zhou, Ke, Jiang, Jiawei, Wang, Chen, Yuan, Binhang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
Cascadia: An Efficient Cascade Serving System for Large Language Models
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026)
von: Liang, Yan, et al.
Veröffentlicht: (2026)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
Parallax: Efficient LLM Inference Service over Decentralized Environment
von: Tong, Chris, et al.
Veröffentlicht: (2025)
von: Tong, Chris, et al.
Veröffentlicht: (2025)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
von: He, Guoliang, et al.
Veröffentlicht: (2025)
von: He, Guoliang, et al.
Veröffentlicht: (2025)
FATE: Future-State-Aware Scheduling for Heterogeneous LLM Workflows
von: Huang, Zirui, et al.
Veröffentlicht: (2026)
von: Huang, Zirui, et al.
Veröffentlicht: (2026)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
von: He, Xuan, et al.
Veröffentlicht: (2025)
von: He, Xuan, et al.
Veröffentlicht: (2025)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
von: Dai, Yinwei, et al.
Veröffentlicht: (2025)
von: Dai, Yinwei, et al.
Veröffentlicht: (2025)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
von: Zhang, Han, et al.
Veröffentlicht: (2026)
von: Zhang, Han, et al.
Veröffentlicht: (2026)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
von: Li, Pengbo, et al.
Veröffentlicht: (2026)
von: Li, Pengbo, et al.
Veröffentlicht: (2026)
Memory-aware Adaptive Scheduling of Scientific Workflows on Heterogeneous Architectures
von: Kulagina, Svetlana, et al.
Veröffentlicht: (2025)
von: Kulagina, Svetlana, et al.
Veröffentlicht: (2025)
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
WOW: Workflow-Aware Data Movement and Task Scheduling for Dynamic Scientific Workflows
von: Lehmann, Fabian, et al.
Veröffentlicht: (2025)
von: Lehmann, Fabian, et al.
Veröffentlicht: (2025)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
von: Mo, Zizhao, et al.
Veröffentlicht: (2025)
von: Mo, Zizhao, et al.
Veröffentlicht: (2025)
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
von: Wei, Xinming, et al.
Veröffentlicht: (2025)
von: Wei, Xinming, et al.
Veröffentlicht: (2025)
Carbon-Aware Mapping and Scheduling for Deadline-Constrained Workflows
von: Schweisgut, Dominik, et al.
Veröffentlicht: (2026)
von: Schweisgut, Dominik, et al.
Veröffentlicht: (2026)
Efficient Probabilistic Workflow Scheduling for IaaS Clouds
von: Russo, Gabriele Russo, et al.
Veröffentlicht: (2024)
von: Russo, Gabriele Russo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2025) -
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
von: Jiang, Youhe, et al.
Veröffentlicht: (2026) -
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2023) -
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025) -
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)