OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, June, Xu, Neal, Huang, Gragas, Zhou, Bok, Liu, Stephen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Efficient and Scalable Distributed Vector Search with RDMA
por: Zhi, Xiangyu, et al.
Publicado: (2025)
por: Zhi, Xiangyu, et al.
Publicado: (2025)
ALock: Asymmetric Lock Primitive for RDMA Systems
por: Baran, Amanda, et al.
Publicado: (2024)
por: Baran, Amanda, et al.
Publicado: (2024)
fabric-lib: RDMA Point-to-Point Communication for LLM Systems
por: Licker, Nandor, et al.
Publicado: (2025)
por: Licker, Nandor, et al.
Publicado: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
por: Brock, Benjamin, et al.
Publicado: (2023)
por: Brock, Benjamin, et al.
Publicado: (2023)
The Semantic Arrow of Time, Part III: RDMA and the Completion Fallacy
por: Borrill, Paul
Publicado: (2026)
por: Borrill, Paul
Publicado: (2026)
RHAPSODY: Execution of Hybrid AI-HPC Workflows at Scale
por: Alsaadi, Aymen, et al.
Publicado: (2025)
por: Alsaadi, Aymen, et al.
Publicado: (2025)
Reducing the Impact of I/O Contention in Numerical Weather Prediction Workflows at Scale Using DAOS
por: Manubens, Nicolau, et al.
Publicado: (2024)
por: Manubens, Nicolau, et al.
Publicado: (2024)
Unleashing the Power of Tree-of-Thoughts for Edge-Enabled AIGC Service Provisioning
por: Liu, Zhang, et al.
Publicado: (2026)
por: Liu, Zhang, et al.
Publicado: (2026)
Steering a Fleet: Adaptation for Large-Scale, Workflow-Based Experiments
por: Pruyne, Jim, et al.
Publicado: (2024)
por: Pruyne, Jim, et al.
Publicado: (2024)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
por: Chen, Jiu, et al.
Publicado: (2026)
por: Chen, Jiu, et al.
Publicado: (2026)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
por: Liu, Xing, et al.
Publicado: (2025)
por: Liu, Xing, et al.
Publicado: (2025)
Batch Denoising for AIGC Service Provisioning in Wireless Edge Networks
por: Xu, Jinghang, et al.
Publicado: (2025)
por: Xu, Jinghang, et al.
Publicado: (2025)
FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
por: Guo, Xianwen, et al.
Publicado: (2025)
por: Guo, Xianwen, et al.
Publicado: (2025)
Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
por: Stavrinides, Georgios L., et al.
Publicado: (2025)
por: Stavrinides, Georgios L., et al.
Publicado: (2025)
Distributed Inference Performance Optimization for LLMs on CPUs
por: He, Pujiang, et al.
Publicado: (2024)
por: He, Pujiang, et al.
Publicado: (2024)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
por: Yu, Minchen, et al.
Publicado: (2025)
por: Yu, Minchen, et al.
Publicado: (2025)
CoGenT: A Content-oriented Generative-hit Framework for Content Delivery Networks
por: Wang, Peng, et al.
Publicado: (2024)
por: Wang, Peng, et al.
Publicado: (2024)
Scaling on Frontier: Uncertainty Quantification Workflow Applications using ExaWorks to Enable Full System Utilization
por: Titov, Mikhail, et al.
Publicado: (2024)
por: Titov, Mikhail, et al.
Publicado: (2024)
A Terminology for Scientific Workflow Systems
por: Suter, Frédéric, et al.
Publicado: (2025)
por: Suter, Frédéric, et al.
Publicado: (2025)
Reimagining RDMA Through the Lens of ML
por: Warraich, Ertza, et al.
Publicado: (2025)
por: Warraich, Ertza, et al.
Publicado: (2025)
Characterizing Communication Patterns in Distributed Large Language Model Inference
por: Xu, Lang, et al.
Publicado: (2025)
por: Xu, Lang, et al.
Publicado: (2025)
AI-coupled HPC Workflow Applications, Middleware and Performance
por: Brewer, Wes, et al.
Publicado: (2024)
por: Brewer, Wes, et al.
Publicado: (2024)
iDDS: Intelligent Distributed Dispatch and Scheduling for Workflow Orchestration
por: Guan, Wen, et al.
Publicado: (2025)
por: Guan, Wen, et al.
Publicado: (2025)
Handling of Memory Page Faults during Virtual-Address RDMA
por: Psistakis, Antonis
Publicado: (2025)
por: Psistakis, Antonis
Publicado: (2025)
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
por: Razavi, Kamran, et al.
Publicado: (2024)
por: Razavi, Kamran, et al.
Publicado: (2024)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
por: Jiang, Youhe, et al.
Publicado: (2023)
por: Jiang, Youhe, et al.
Publicado: (2023)
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
por: Warraich, Ertza, et al.
Publicado: (2025)
por: Warraich, Ertza, et al.
Publicado: (2025)
Leveraging Core and Uncore Frequency Scaling for Power-Efficient Serverless Workflows
por: Tzenetopoulos, Achilleas, et al.
Publicado: (2024)
por: Tzenetopoulos, Achilleas, et al.
Publicado: (2024)
AIGC-assisted Federated Learning for Vehicular Edge Intelligence: Vehicle Selection, Resource Allocation and Model Augmentation
por: Qiang, Xianke, et al.
Publicado: (2025)
por: Qiang, Xianke, et al.
Publicado: (2025)
EcoServe: Designing Carbon-Aware AI Inference Systems
por: Li, Yueying, et al.
Publicado: (2025)
por: Li, Yueying, et al.
Publicado: (2025)
An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
por: Hakim, Sheikh Azizul, et al.
Publicado: (2025)
por: Hakim, Sheikh Azizul, et al.
Publicado: (2025)
Mapping Large Memory-constrained Workflows onto Heterogeneous Platforms
por: Kulagina, Svetlana, et al.
Publicado: (2024)
por: Kulagina, Svetlana, et al.
Publicado: (2024)
FedRDMA: Communication-Efficient Cross-Silo Federated LLM via Chunked RDMA Transmission
por: Zhang, Zeling, et al.
Publicado: (2024)
por: Zhang, Zeling, et al.
Publicado: (2024)
In-Transit Data Transport Strategies for Coupled AI-Simulation Workflow Patterns
por: Tummalapalli, Harikrishna, et al.
Publicado: (2025)
por: Tummalapalli, Harikrishna, et al.
Publicado: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
por: Huang, Jinqi, et al.
Publicado: (2025)
por: Huang, Jinqi, et al.
Publicado: (2025)
EAT: QoS-Aware Edge-Collaborative AIGC Task Scheduling via Attention-Guided Diffusion Reinforcement Learning
por: Xu, Zhifei, et al.
Publicado: (2025)
por: Xu, Zhifei, et al.
Publicado: (2025)
Large-Scale LLM Inference with Heterogeneous Workloads: Prefill-Decode Contention and Asymptotically Optimal Control
por: Lin, Ruihan, et al.
Publicado: (2026)
por: Lin, Ruihan, et al.
Publicado: (2026)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
por: Ahmad, Sohaib, et al.
Publicado: (2024)
por: Ahmad, Sohaib, et al.
Publicado: (2024)
Efficiently Reproducing Distributed Workflows in Notebook-based Systems
por: Azaz, Talha, et al.
Publicado: (2026)
por: Azaz, Talha, et al.
Publicado: (2026)
Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
por: Li, Rui, et al.
Publicado: (2025)
por: Li, Rui, et al.
Publicado: (2025)
Ejemplares similares
-
Towards Efficient and Scalable Distributed Vector Search with RDMA
por: Zhi, Xiangyu, et al.
Publicado: (2025) -
ALock: Asymmetric Lock Primitive for RDMA Systems
por: Baran, Amanda, et al.
Publicado: (2024) -
fabric-lib: RDMA Point-to-Point Communication for LLM Systems
por: Licker, Nandor, et al.
Publicado: (2025) -
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
por: Brock, Benjamin, et al.
Publicado: (2023) -
The Semantic Arrow of Time, Part III: RDMA and the Completion Fallacy
por: Borrill, Paul
Publicado: (2026)