Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
Fuente:
arXiv
Salvato in:
| Autori principali: | Wei, Jinhui, Huang, Ye, Zhou, Yuhui, Jiang, Jiazhi, Du, Jiangsu, Lu, Yutong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
di: Zhang, Hongbin, et al.
Pubblicazione: (2026)
di: Zhang, Hongbin, et al.
Pubblicazione: (2026)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
di: Bai, Fengyao, et al.
Pubblicazione: (2026)
di: Bai, Fengyao, et al.
Pubblicazione: (2026)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
di: Zhang, Yida, et al.
Pubblicazione: (2026)
di: Zhang, Yida, et al.
Pubblicazione: (2026)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
di: Tran, Phuong, et al.
Pubblicazione: (2025)
di: Tran, Phuong, et al.
Pubblicazione: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
di: Han, Yunhe, et al.
Pubblicazione: (2026)
di: Han, Yunhe, et al.
Pubblicazione: (2026)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
di: Wang, Zhibin, et al.
Pubblicazione: (2025)
di: Wang, Zhibin, et al.
Pubblicazione: (2025)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
di: Li, Yuchen, et al.
Pubblicazione: (2026)
di: Li, Yuchen, et al.
Pubblicazione: (2026)
Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
di: Ye, Shengyuan, et al.
Pubblicazione: (2024)
di: Ye, Shengyuan, et al.
Pubblicazione: (2024)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
di: Chen, Wenyan, et al.
Pubblicazione: (2026)
di: Chen, Wenyan, et al.
Pubblicazione: (2026)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
di: Dai, Yuntao, et al.
Pubblicazione: (2026)
di: Dai, Yuntao, et al.
Pubblicazione: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
di: Liu, Xing, et al.
Pubblicazione: (2025)
di: Liu, Xing, et al.
Pubblicazione: (2025)
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
di: Zheng, Ce, et al.
Pubblicazione: (2026)
di: Zheng, Ce, et al.
Pubblicazione: (2026)
PolyKAN: Efficient Fused GPU Operators for Polynomial Kolmogorov-Arnold Network Variants
di: Yu, Mingkun, et al.
Pubblicazione: (2025)
di: Yu, Mingkun, et al.
Pubblicazione: (2025)
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
di: Venkatesha, Yeshwanth, et al.
Pubblicazione: (2025)
di: Venkatesha, Yeshwanth, et al.
Pubblicazione: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
di: Ramachandran, Arun, et al.
Pubblicazione: (2025)
di: Ramachandran, Arun, et al.
Pubblicazione: (2025)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
di: Song, Jingwei, et al.
Pubblicazione: (2025)
di: Song, Jingwei, et al.
Pubblicazione: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
di: Gao, Luyao, et al.
Pubblicazione: (2025)
di: Gao, Luyao, et al.
Pubblicazione: (2025)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
di: Zhang, Ziyi, et al.
Pubblicazione: (2025)
di: Zhang, Ziyi, et al.
Pubblicazione: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
di: Chen, Liangkun, et al.
Pubblicazione: (2025)
di: Chen, Liangkun, et al.
Pubblicazione: (2025)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
di: Shen, Yuhao, et al.
Pubblicazione: (2025)
di: Shen, Yuhao, et al.
Pubblicazione: (2025)
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
di: Cheng, Rongxin, et al.
Pubblicazione: (2025)
di: Cheng, Rongxin, et al.
Pubblicazione: (2025)
Tasking framework for Adaptive Speculative Parallel Mesh Generation
di: Tsolakis, Christos, et al.
Pubblicazione: (2024)
di: Tsolakis, Christos, et al.
Pubblicazione: (2024)
Distributed Semi-Speculative Parallel Anisotropic Mesh Adaptation
di: Garner, Kevin, et al.
Pubblicazione: (2026)
di: Garner, Kevin, et al.
Pubblicazione: (2026)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
di: Ma, Tenghui, et al.
Pubblicazione: (2026)
di: Ma, Tenghui, et al.
Pubblicazione: (2026)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
di: Zhang, Mingjin, et al.
Pubblicazione: (2024)
di: Zhang, Mingjin, et al.
Pubblicazione: (2024)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
di: Lu, Yao, et al.
Pubblicazione: (2026)
di: Lu, Yao, et al.
Pubblicazione: (2026)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
di: Xu, Wendong, et al.
Pubblicazione: (2025)
di: Xu, Wendong, et al.
Pubblicazione: (2025)
SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative Models
di: Chen, Fahao, et al.
Pubblicazione: (2025)
di: Chen, Fahao, et al.
Pubblicazione: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
di: Chen, Haoyu, et al.
Pubblicazione: (2025)
di: Chen, Haoyu, et al.
Pubblicazione: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
di: Arya, Mayank, et al.
Pubblicazione: (2025)
di: Arya, Mayank, et al.
Pubblicazione: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
di: Rajashekar, Kolichala, et al.
Pubblicazione: (2025)
di: Rajashekar, Kolichala, et al.
Pubblicazione: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
di: Zhao, Alan, et al.
Pubblicazione: (2026)
di: Zhao, Alan, et al.
Pubblicazione: (2026)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
di: Khoshsirat, Aria, et al.
Pubblicazione: (2024)
di: Khoshsirat, Aria, et al.
Pubblicazione: (2024)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
di: Wang, Chong, et al.
Pubblicazione: (2026)
di: Wang, Chong, et al.
Pubblicazione: (2026)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
di: Li, Zixuan, et al.
Pubblicazione: (2024)
di: Li, Zixuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
di: Zhang, Hongbin, et al.
Pubblicazione: (2025) -
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
di: Zhang, Hongbin, et al.
Pubblicazione: (2026) -
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
di: Guo, Tianyu, et al.
Pubblicazione: (2025) -
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
di: Du, Jiangsu, et al.
Pubblicazione: (2025) -
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
di: Bai, Fengyao, et al.
Pubblicazione: (2026)