Guardado en:
| Autores principales: | Liu, Zhibang, Xu, Chaonong, Liu, Zhizhuo, Huang, Lekai, Wei, Jiachen, Li, Chao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2409.07693 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
por: Liu, Zhibang, et al.
Publicado: (2025)
por: Liu, Zhibang, et al.
Publicado: (2025)
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
por: Yang, Xiang, et al.
Publicado: (2022)
por: Yang, Xiang, et al.
Publicado: (2022)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
por: Liu, Xing, et al.
Publicado: (2025)
por: Liu, Xing, et al.
Publicado: (2025)
Learning the Optimal Path and DNN Partition for Collaborative Edge Inference
por: Huang, Yin, et al.
Publicado: (2024)
por: Huang, Yin, et al.
Publicado: (2024)
Where to Split? A Pareto-Front Analysis of DNN Partitioning for Edge Inference
por: Masud, Adiba, et al.
Publicado: (2026)
por: Masud, Adiba, et al.
Publicado: (2026)
Cooperative Gradient Coding
por: Weng, Shudi, et al.
Publicado: (2025)
por: Weng, Shudi, et al.
Publicado: (2025)
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
por: Duan, Jiaang, et al.
Publicado: (2024)
por: Duan, Jiaang, et al.
Publicado: (2024)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
por: Chen, Aodong, et al.
Publicado: (2023)
por: Chen, Aodong, et al.
Publicado: (2023)
Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Optimization
por: Gao, Luyao, et al.
Publicado: (2024)
por: Gao, Luyao, et al.
Publicado: (2024)
Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative Caching
por: Liang, Wenyi, et al.
Publicado: (2024)
por: Liang, Wenyi, et al.
Publicado: (2024)
WindGP: Efficient Graph Partitioning on Heterogenous Machines
por: Zeng, Li, et al.
Publicado: (2024)
por: Zeng, Li, et al.
Publicado: (2024)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
por: Liu, Xing, et al.
Publicado: (2025)
por: Liu, Xing, et al.
Publicado: (2025)
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
por: Chen, June, et al.
Publicado: (2026)
por: Chen, June, et al.
Publicado: (2026)
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
por: Zheng, Zihao, et al.
Publicado: (2026)
por: Zheng, Zihao, et al.
Publicado: (2026)
SparseMap: Loop Mapping for Sparse CNNs on Streaming Coarse-grained Reconfigurable Array
por: Ni, Xiaobing, et al.
Publicado: (2024)
por: Ni, Xiaobing, et al.
Publicado: (2024)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
por: Tang, Xinru, et al.
Publicado: (2025)
por: Tang, Xinru, et al.
Publicado: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
por: Chen, Jiabin, et al.
Publicado: (2024)
por: Chen, Jiabin, et al.
Publicado: (2024)
Partition Detection in Byzantine Networks
por: Bromberg, Yérom-David, et al.
Publicado: (2024)
por: Bromberg, Yérom-David, et al.
Publicado: (2024)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
por: Hu, Cunchen, et al.
Publicado: (2024)
por: Hu, Cunchen, et al.
Publicado: (2024)
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
por: Saba, Issa, et al.
Publicado: (2024)
por: Saba, Issa, et al.
Publicado: (2024)
SLO-Aware Scheduling for Large Language Model Inferences
por: Huang, Jinqi, et al.
Publicado: (2025)
por: Huang, Jinqi, et al.
Publicado: (2025)
Persistent and Partitioned MPI for Stencil Communication
por: Collom, Gerald, et al.
Publicado: (2025)
por: Collom, Gerald, et al.
Publicado: (2025)
Incidence Constraints in Hypergraph Partitioning on GPU
por: Ronzani, Marco, et al.
Publicado: (2026)
por: Ronzani, Marco, et al.
Publicado: (2026)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
por: Zhao, Alan, et al.
Publicado: (2026)
por: Zhao, Alan, et al.
Publicado: (2026)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
por: Gao, Wei, et al.
Publicado: (2026)
por: Gao, Wei, et al.
Publicado: (2026)
Large Language Model Partitioning for Low-Latency Inference at the Edge
por: Kafetzis, Dimitrios, et al.
Publicado: (2025)
por: Kafetzis, Dimitrios, et al.
Publicado: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
por: Gao, Luyao, et al.
Publicado: (2025)
por: Gao, Luyao, et al.
Publicado: (2025)
YUHENG-OS: A Cloud-Native Space Cluster Operating System
por: Zhang, Jin, et al.
Publicado: (2026)
por: Zhang, Jin, et al.
Publicado: (2026)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
por: Lin, Haoran, et al.
Publicado: (2025)
por: Lin, Haoran, et al.
Publicado: (2025)
FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
por: Chen, Huamin, et al.
Publicado: (2026)
por: Chen, Huamin, et al.
Publicado: (2026)
A Distributed Partitioning Software and its Applications
por: Sasidharan, Aparna
Publicado: (2025)
por: Sasidharan, Aparna
Publicado: (2025)
Deterministic Parallel High-Quality Hypergraph Partitioning
por: Krause, Robert, et al.
Publicado: (2025)
por: Krause, Robert, et al.
Publicado: (2025)
PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
por: Zhao, Zhixin, et al.
Publicado: (2026)
por: Zhao, Zhixin, et al.
Publicado: (2026)
FedQuad: Adaptive Layer-wise LoRA Deployment and Activation Quantization for Federated Fine-Tuning
por: Li, Rukuo, et al.
Publicado: (2025)
por: Li, Rukuo, et al.
Publicado: (2025)
inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference
por: Chen, Huamin, et al.
Publicado: (2026)
por: Chen, Huamin, et al.
Publicado: (2026)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
por: Zhang, Ziyang, et al.
Publicado: (2025)
por: Zhang, Ziyang, et al.
Publicado: (2025)
KV Cache Compression for Inference Efficiency in LLMs: A Review
por: Liu, Yanyu, et al.
Publicado: (2025)
por: Liu, Yanyu, et al.
Publicado: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
por: Wei, Jinhui, et al.
Publicado: (2025)
por: Wei, Jinhui, et al.
Publicado: (2025)
Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems
por: Kreß, Fabian, et al.
Publicado: (2024)
por: Kreß, Fabian, et al.
Publicado: (2024)
DAWN: Matrix Operation-Optimized Algorithm for Shortest Paths Problem on Unweighted Graphs
por: Feng, Yelai, et al.
Publicado: (2022)
por: Feng, Yelai, et al.
Publicado: (2022)
Ejemplares similares
-
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
por: Liu, Zhibang, et al.
Publicado: (2025) -
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
por: Yang, Xiang, et al.
Publicado: (2022) -
CoCoI: Distributed Coded Inference System for Straggler Mitigation
por: Liu, Xing, et al.
Publicado: (2025) -
Learning the Optimal Path and DNN Partition for Collaborative Edge Inference
por: Huang, Yin, et al.
Publicado: (2024) -
Where to Split? A Pareto-Front Analysis of DNN Partitioning for Edge Inference
por: Masud, Adiba, et al.
Publicado: (2026)