Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Haoyu, Li, Xue, Qian, Kun, Guan, Yu, Zhao, Jin, Wang, Xin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
di: Xu, Wendong, et al.
Pubblicazione: (2025)
di: Xu, Wendong, et al.
Pubblicazione: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
di: Zhao, Yihao, et al.
Pubblicazione: (2025)
di: Zhao, Yihao, et al.
Pubblicazione: (2025)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
di: Srivatsa, Vikranth, et al.
Pubblicazione: (2026)
di: Srivatsa, Vikranth, et al.
Pubblicazione: (2026)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
di: Da, Wei, et al.
Pubblicazione: (2026)
di: Da, Wei, et al.
Pubblicazione: (2026)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
di: Shyam, Vasu, et al.
Pubblicazione: (2026)
di: Shyam, Vasu, et al.
Pubblicazione: (2026)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
di: Ai, Xin, et al.
Pubblicazione: (2024)
di: Ai, Xin, et al.
Pubblicazione: (2024)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
di: Zhao, Alan, et al.
Pubblicazione: (2026)
di: Zhao, Alan, et al.
Pubblicazione: (2026)
Synergistic Tensor and Pipeline Parallelism
di: Qi, Mengshi, et al.
Pubblicazione: (2025)
di: Qi, Mengshi, et al.
Pubblicazione: (2025)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
di: Xiang, Yuxing, et al.
Pubblicazione: (2025)
di: Xiang, Yuxing, et al.
Pubblicazione: (2025)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
di: Tang, Ding, et al.
Pubblicazione: (2024)
di: Tang, Ding, et al.
Pubblicazione: (2024)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
di: Wang, Zhigang, et al.
Pubblicazione: (2024)
di: Wang, Zhigang, et al.
Pubblicazione: (2024)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
di: Lin, Haoran, et al.
Pubblicazione: (2025)
di: Lin, Haoran, et al.
Pubblicazione: (2025)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
di: Liu, Ruitao, et al.
Pubblicazione: (2026)
di: Liu, Ruitao, et al.
Pubblicazione: (2026)
Parallax: Efficient LLM Inference Service over Decentralized Environment
di: Tong, Chris, et al.
Pubblicazione: (2025)
di: Tong, Chris, et al.
Pubblicazione: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
di: He, Yongchao, et al.
Pubblicazione: (2025)
di: He, Yongchao, et al.
Pubblicazione: (2025)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
di: Li, Cong, et al.
Pubblicazione: (2025)
di: Li, Cong, et al.
Pubblicazione: (2025)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
di: Fan, Jiakun, et al.
Pubblicazione: (2025)
di: Fan, Jiakun, et al.
Pubblicazione: (2025)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
di: Duan, Jiaang, et al.
Pubblicazione: (2024)
di: Duan, Jiaang, et al.
Pubblicazione: (2024)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
di: Liu, Zihan, et al.
Pubblicazione: (2025)
di: Liu, Zihan, et al.
Pubblicazione: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
di: Wang, Weiye, et al.
Pubblicazione: (2026)
di: Wang, Weiye, et al.
Pubblicazione: (2026)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
di: Wei, Jinhui, et al.
Pubblicazione: (2025)
di: Wei, Jinhui, et al.
Pubblicazione: (2025)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
di: Hidayetoglu, Mert, et al.
Pubblicazione: (2025)
di: Hidayetoglu, Mert, et al.
Pubblicazione: (2025)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
di: Liu, Zhibang, et al.
Pubblicazione: (2025)
di: Liu, Zhibang, et al.
Pubblicazione: (2025)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
di: Liu, Man, et al.
Pubblicazione: (2026)
di: Liu, Man, et al.
Pubblicazione: (2026)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
di: Ma, Bin, et al.
Pubblicazione: (2026)
di: Ma, Bin, et al.
Pubblicazione: (2026)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
di: Luo, Jiajun, et al.
Pubblicazione: (2024)
di: Luo, Jiajun, et al.
Pubblicazione: (2024)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
di: Kong, Jie, et al.
Pubblicazione: (2026)
di: Kong, Jie, et al.
Pubblicazione: (2026)
A New Execution Model and Executor for Adaptively Optimizing the Performance of Parallel Algorithms Using HPX Runtime System
di: Mohammadiporshokooh, Karame, et al.
Pubblicazione: (2025)
di: Mohammadiporshokooh, Karame, et al.
Pubblicazione: (2025)
GTaP: A GPU-Resident Fork-Join Task-Parallel Runtime with a Pragma-Based Interface
di: Maeda, Yuki, et al.
Pubblicazione: (2026)
di: Maeda, Yuki, et al.
Pubblicazione: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
Minimizing Communication for Parallel Symmetric Tensor Times Same Vector Computation
di: Daas, Hussam Al, et al.
Pubblicazione: (2025)
di: Daas, Hussam Al, et al.
Pubblicazione: (2025)
Fail-Closed Lowering of Resident KV Claims onto LLM Serving Runtimes
di: Stepanek, Lukas
Pubblicazione: (2026)
di: Stepanek, Lukas
Pubblicazione: (2026)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
di: Zhao, Long, et al.
Pubblicazione: (2026)
di: Zhao, Long, et al.
Pubblicazione: (2026)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
di: Liu, Di, et al.
Pubblicazione: (2026)
di: Liu, Di, et al.
Pubblicazione: (2026)
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
di: Liu, Ziming, et al.
Pubblicazione: (2024)
di: Liu, Ziming, et al.
Pubblicazione: (2024)
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
di: Wu, Yongtong, et al.
Pubblicazione: (2026)
di: Wu, Yongtong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
di: Xu, Wendong, et al.
Pubblicazione: (2025) -
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
di: Zhao, Yihao, et al.
Pubblicazione: (2025) -
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
di: Srivatsa, Vikranth, et al.
Pubblicazione: (2026) -
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
di: Da, Wei, et al.
Pubblicazione: (2026) -
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
di: Shyam, Vasu, et al.
Pubblicazione: (2026)