Gespeichert in:
| Hauptverfasser: | Wang, Peng, Liu, Yu, Liu, Ziqi, Wang, Ming-Yang, Liu, Ke, Zhou, Ke, Huang, Zhihai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2402.16262 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
von: Chen, June, et al.
Veröffentlicht: (2026)
von: Chen, June, et al.
Veröffentlicht: (2026)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
Content-Oblivious Leader Election in 2-Edge-Connected Networks
von: Chalopin, Jérémie, et al.
Veröffentlicht: (2025)
von: Chalopin, Jérémie, et al.
Veröffentlicht: (2025)
Big Data-Driven Fraud Detection Using Machine Learning and Real-Time Stream Processing
von: Liu, Chen, et al.
Veröffentlicht: (2025)
von: Liu, Chen, et al.
Veröffentlicht: (2025)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
von: Wu, Siyu, et al.
Veröffentlicht: (2025)
von: Wu, Siyu, et al.
Veröffentlicht: (2025)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
von: Kong, Jie, et al.
Veröffentlicht: (2026)
von: Kong, Jie, et al.
Veröffentlicht: (2026)
Efficient Counting and Simulation in Content-Oblivious Rings
von: Chalopin, Jérémie, et al.
Veröffentlicht: (2026)
von: Chalopin, Jérémie, et al.
Veröffentlicht: (2026)
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
Integrated Sensing, Communication, and Computing: An Information-oriented Resource Transaction Mechanism
von: Chen, Ning, et al.
Veröffentlicht: (2024)
von: Chen, Ning, et al.
Veröffentlicht: (2024)
A Survey on Adversarial Contention Resolution
von: Banicescu, Ioana, et al.
Veröffentlicht: (2024)
von: Banicescu, Ioana, et al.
Veröffentlicht: (2024)
Softening the Impact of Collisions in Contention Resolution
von: Biswas, Umesh, et al.
Veröffentlicht: (2024)
von: Biswas, Umesh, et al.
Veröffentlicht: (2024)
Beyond 2-Edge-Connectivity: Algorithms and Impossibility for Content-Oblivious Leader Election
von: Chang, Yi-Jun, et al.
Veröffentlicht: (2025)
von: Chang, Yi-Jun, et al.
Veröffentlicht: (2025)
Non-Uniform Content-Oblivious Leader Election on Oriented Asynchronous Rings
von: Chalopin, Jérémie, et al.
Veröffentlicht: (2025)
von: Chalopin, Jérémie, et al.
Veröffentlicht: (2025)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
TIDAL: Recovering Temporal Phase for Cloud Block Storage Placement from LLM-Derived Semantics
von: Tan, Difan, et al.
Veröffentlicht: (2026)
von: Tan, Difan, et al.
Veröffentlicht: (2026)
QoE-oriented Dependent Task Scheduling under Multi-dimensional QoS Constraints over Distributed Networks
von: Fan, Xuwei, et al.
Veröffentlicht: (2023)
von: Fan, Xuwei, et al.
Veröffentlicht: (2023)
Graph-Structured Deep Learning Framework for Multi-task Contention Identification with High-dimensional Metrics
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
A Contention-Free Model for Converged Kubernetes on HPC
von: Sochat, Vanessa, et al.
Veröffentlicht: (2024)
von: Sochat, Vanessa, et al.
Veröffentlicht: (2024)
Warp-STAR: High-performance, Differentiable GPU-Accelerated Static Timing Analysis through Warp-oriented Parallel Orchestration
von: Huang, En-Ming, et al.
Veröffentlicht: (2026)
von: Huang, En-Ming, et al.
Veröffentlicht: (2026)
FedHC: A Hierarchical Clustered Federated Learning Framework for Satellite Networks
von: Liu, Zhuocheng, et al.
Veröffentlicht: (2025)
von: Liu, Zhuocheng, et al.
Veröffentlicht: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
Accelerating Microswimmer Simulations via a Heterogeneous Pipelined Parallel-in-Time Framework
von: Huang, Ruixiang, et al.
Veröffentlicht: (2026)
von: Huang, Ruixiang, et al.
Veröffentlicht: (2026)
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
von: Zhao, Zhixin, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixin, et al.
Veröffentlicht: (2024)
FedCache: A Knowledge Cache-driven Federated Learning Architecture for Personalized Edge Intelligence
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2023)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
von: Wu, Yu, et al.
Veröffentlicht: (2025)
von: Wu, Yu, et al.
Veröffentlicht: (2025)
A Thorough Investigation of Content-Defined Chunking Algorithms for Data Deduplication
von: Gregoriadis, Marcel, et al.
Veröffentlicht: (2024)
von: Gregoriadis, Marcel, et al.
Veröffentlicht: (2024)
AdaBridge: Dynamic Data and Computation Reuse for Efficient Multi-task DNN Co-evolution in Edge Systems
von: Wang, Lehao, et al.
Veröffentlicht: (2024)
von: Wang, Lehao, et al.
Veröffentlicht: (2024)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025)
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
Accelerating the Delivery of Data Services over Uncertain Mobile Crowdsensing Networks
von: Liwang, Minghui, et al.
Veröffentlicht: (2022)
von: Liwang, Minghui, et al.
Veröffentlicht: (2022)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
ConChain: A Scheme for Contention-free and Attack Resilient BlockChain
von: Bappy, Faisal Haque, et al.
Veröffentlicht: (2023)
von: Bappy, Faisal Haque, et al.
Veröffentlicht: (2023)
BandPilot: Towards Performance- and Contention-Aware GPU Dispatching in AI Clusters
von: Zhang, Kunming, et al.
Veröffentlicht: (2025)
von: Zhang, Kunming, et al.
Veröffentlicht: (2025)
Contention Resolution, With and Without a Global Clock
von: Cai, Zixi, et al.
Veröffentlicht: (2026)
von: Cai, Zixi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
von: Chen, June, et al.
Veröffentlicht: (2026) -
CoCoI: Distributed Coded Inference System for Straggler Mitigation
von: Liu, Xing, et al.
Veröffentlicht: (2025) -
Content-Oblivious Leader Election in 2-Edge-Connected Networks
von: Chalopin, Jérémie, et al.
Veröffentlicht: (2025) -
Big Data-Driven Fraud Detection Using Machine Learning and Real-Time Stream Processing
von: Liu, Chen, et al.
Veröffentlicht: (2025) -
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)