MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Lei, Chen, Mouxiang, Cao, Ruisheng, Chen, Jiawei, Zhou, Fan, Xu, Yiheng, Yang, Jiaxi, Ma, Zeyao, Chen, Liang, Luo, Changwei, Zhang, Kai, Yan, Fan, Shum, KaShun, Zhang, Jiajun, Cui, Zeyu, Hu, Feng, Lin, Junyang, Hui, Binyuan, Yang, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SWE-RM: Execution-free Feedback For Software Engineering Agents
von: Shum, KaShun, et al.
Veröffentlicht: (2025)
von: Shum, KaShun, et al.
Veröffentlicht: (2025)
SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Manner
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data
von: Shum, KaShun, et al.
Veröffentlicht: (2023)
von: Shum, KaShun, et al.
Veröffentlicht: (2023)
Qwen3-Coder-Next Technical Report
von: Cao, Ruisheng, et al.
Veröffentlicht: (2026)
von: Cao, Ruisheng, et al.
Veröffentlicht: (2026)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
von: Lu, Yishun, et al.
Veröffentlicht: (2026)
von: Lu, Yishun, et al.
Veröffentlicht: (2026)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
von: Li, Maoliang, et al.
Veröffentlicht: (2026)
von: Li, Maoliang, et al.
Veröffentlicht: (2026)
Scaling Agentic Verifier for Competitive Coding
von: Ma, Zeyao, et al.
Veröffentlicht: (2026)
von: Ma, Zeyao, et al.
Veröffentlicht: (2026)
Parallel Scaling Law for Language Models
von: Chen, Mouxiang, et al.
Veröffentlicht: (2025)
von: Chen, Mouxiang, et al.
Veröffentlicht: (2025)
Agentic AI-Driven UAV Network Deployment: An LLM-Enhanced Exact Potential Game Approach
von: Tang, Xin, et al.
Veröffentlicht: (2026)
von: Tang, Xin, et al.
Veröffentlicht: (2026)
MegaFlow: Zero-Shot Large Displacement Optical Flow
von: Zhang, Dingxi, et al.
Veröffentlicht: (2026)
von: Zhang, Dingxi, et al.
Veröffentlicht: (2026)
Hurry: Dynamic Collaborative Framework For Low-orbit Mega-Constellation Data Downloading
von: Luo, Handong, et al.
Veröffentlicht: (2024)
von: Luo, Handong, et al.
Veröffentlicht: (2024)
SWE-Universe: Scale Real-World Verifiable Environments to Millions
von: Chen, Mouxiang, et al.
Veröffentlicht: (2026)
von: Chen, Mouxiang, et al.
Veröffentlicht: (2026)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
von: Biswas, Anish, et al.
Veröffentlicht: (2026)
von: Biswas, Anish, et al.
Veröffentlicht: (2026)
FusionRCG: Orchestrating Recursive Computation Graphs across GPU Memory Hierarchies
von: Zhang, Yihong, et al.
Veröffentlicht: (2026)
von: Zhang, Yihong, et al.
Veröffentlicht: (2026)
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR
von: Zhang, Yiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Yiqi, et al.
Veröffentlicht: (2026)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
Optimizing Federated Learning in the Era of LLMs: Message Quantization and Streaming
von: Xu, Ziyue, et al.
Veröffentlicht: (2025)
von: Xu, Ziyue, et al.
Veröffentlicht: (2025)
Boosting Scientific Error-Bounded Lossy Compression through Optimized Synergistic Lossy-Lossless Orchestration
von: Wu, Shixun, et al.
Veröffentlicht: (2025)
von: Wu, Shixun, et al.
Veröffentlicht: (2025)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
von: Zhan, Huiyou, et al.
Veröffentlicht: (2025)
von: Zhan, Huiyou, et al.
Veröffentlicht: (2025)
AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
von: Tokal, Shiva Sai Krishna Anand, et al.
Veröffentlicht: (2025)
von: Tokal, Shiva Sai Krishna Anand, et al.
Veröffentlicht: (2025)
Pier: Efficient Large Language Model pretraining with Relaxed Global Communication
von: Fan, Shuyuan, et al.
Veröffentlicht: (2025)
von: Fan, Shuyuan, et al.
Veröffentlicht: (2025)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
von: Xuan, Mo, et al.
Veröffentlicht: (2025)
von: Xuan, Mo, et al.
Veröffentlicht: (2025)
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
von: Wu, Yongtong, et al.
Veröffentlicht: (2026)
von: Wu, Yongtong, et al.
Veröffentlicht: (2026)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
Integrated Sensing, Communication, and Computing: An Information-oriented Resource Transaction Mechanism
von: Chen, Ning, et al.
Veröffentlicht: (2024)
von: Chen, Ning, et al.
Veröffentlicht: (2024)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
von: Zhu, Botao, et al.
Veröffentlicht: (2025)
von: Zhu, Botao, et al.
Veröffentlicht: (2025)
Orchestrating Joint Offloading and Scheduling for Low-Latency Edge SLAM
von: Zhang, Yao, et al.
Veröffentlicht: (2025)
von: Zhang, Yao, et al.
Veröffentlicht: (2025)
From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
MegaFold: System-Level Optimizations for Accelerating Protein Structure Prediction Models
von: La, Hoa, et al.
Veröffentlicht: (2025)
von: La, Hoa, et al.
Veröffentlicht: (2025)
Communication-Computation Pipeline Parallel Split Learning over Wireless Edge Networks
von: Liu, Chenyu, et al.
Veröffentlicht: (2025)
von: Liu, Chenyu, et al.
Veröffentlicht: (2025)
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
von: Wei, Xinming, et al.
Veröffentlicht: (2025)
von: Wei, Xinming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SWE-RM: Execution-free Feedback For Software Engineering Agents
von: Shum, KaShun, et al.
Veröffentlicht: (2025) -
SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Manner
von: Zhang, Lei, et al.
Veröffentlicht: (2025) -
Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data
von: Shum, KaShun, et al.
Veröffentlicht: (2023) -
Qwen3-Coder-Next Technical Report
von: Cao, Ruisheng, et al.
Veröffentlicht: (2026) -
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
von: Lu, Yishun, et al.
Veröffentlicht: (2026)