Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Hao Mark, Luk, Wayne, Yiu, Ka Fai Cedric, Li, Rui, Mishchenko, Konstantin, Venieris, Stylianos I., Fan, Hongxiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Progressive Mixed-Precision Decoding for Efficient LLM Inference
di: Chen, Hao Mark, et al.
Pubblicazione: (2024)
di: Chen, Hao Mark, et al.
Pubblicazione: (2024)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
di: Zhang, Zehuan, et al.
Pubblicazione: (2024)
di: Zhang, Zehuan, et al.
Pubblicazione: (2024)
CARIn: Constraint-Aware and Responsive Inference on Heterogeneous Devices for Single- and Multi-DNN Workloads
di: Panopoulos, Ioannis, et al.
Pubblicazione: (2024)
di: Panopoulos, Ioannis, et al.
Pubblicazione: (2024)
Enhancing LLM-based Quantum Code Generation with Multi-Agent Optimization and Quantum Error Correction
di: Campbell, Charlie, et al.
Pubblicazione: (2025)
di: Campbell, Charlie, et al.
Pubblicazione: (2025)
MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
di: Nikolaidis, Sokratis, et al.
Pubblicazione: (2024)
di: Nikolaidis, Sokratis, et al.
Pubblicazione: (2024)
FastTTS: Accelerating Test-Time Scaling for Edge LLM Reasoning
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
di: Chen, Hao Mark, et al.
Pubblicazione: (2026)
di: Chen, Hao Mark, et al.
Pubblicazione: (2026)
MetaML-Pro: Cross-Stage Design Flow Automation for Efficient Deep Learning Acceleration
di: Que, Zhiqiang, et al.
Pubblicazione: (2025)
di: Que, Zhiqiang, et al.
Pubblicazione: (2025)
Accelerating MRI Uncertainty Estimation with Mask-based Bayesian Neural Network
di: Zhang, Zehuan, et al.
Pubblicazione: (2024)
di: Zhang, Zehuan, et al.
Pubblicazione: (2024)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
di: Zhang, Zehuan, et al.
Pubblicazione: (2026)
di: Zhang, Zehuan, et al.
Pubblicazione: (2026)
Enhancing Trustworthiness with Mixed Precision: Benchmarks, Opportunities, and Challenges
di: Lu, Guanxi, et al.
Pubblicazione: (2025)
di: Lu, Guanxi, et al.
Pubblicazione: (2025)
Speculative Decoding with a Speculative Vocabulary
di: Williams, Miles, et al.
Pubblicazione: (2026)
di: Williams, Miles, et al.
Pubblicazione: (2026)
FW-Merging: Scaling Model Merging with Frank-Wolfe Optimization
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
Optimal Control of Switched Systems Governed by Logical Switching Dynamics
di: Zhang, Xiao, et al.
Pubblicazione: (2026)
di: Zhang, Xiao, et al.
Pubblicazione: (2026)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
di: Fan, Ruibo, et al.
Pubblicazione: (2026)
di: Fan, Ruibo, et al.
Pubblicazione: (2026)
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
di: Lu, Guanxi, et al.
Pubblicazione: (2025)
di: Lu, Guanxi, et al.
Pubblicazione: (2025)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
di: Fan, Wang, et al.
Pubblicazione: (2026)
di: Fan, Wang, et al.
Pubblicazione: (2026)
Advancing AI-assisted Hardware Design with Hierarchical Decentralized Training and Personalized Inference-Time Optimization
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
On stochastic control problems with higher-order moments
di: Wang, Yike, et al.
Pubblicazione: (2024)
di: Wang, Yike, et al.
Pubblicazione: (2024)
A-THENA: Early Intrusion Detection for IoT with Time-Aware Hybrid Encoding and Network-Specific Augmentation
di: Panopoulos, Ioannis, et al.
Pubblicazione: (2026)
di: Panopoulos, Ioannis, et al.
Pubblicazione: (2026)
Enhancing Dropout-based Bayesian Neural Networks with Multi-Exit on FPGA
di: Chen, Hao Mark, et al.
Pubblicazione: (2024)
di: Chen, Hao Mark, et al.
Pubblicazione: (2024)
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
di: Yin, Haofei, et al.
Pubblicazione: (2025)
di: Yin, Haofei, et al.
Pubblicazione: (2025)
HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
di: Kwon, Young D., et al.
Pubblicazione: (2025)
di: Kwon, Young D., et al.
Pubblicazione: (2025)
ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
di: Kong, Quan, et al.
Pubblicazione: (2026)
di: Kong, Quan, et al.
Pubblicazione: (2026)
SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations
di: Wang, Zhican, et al.
Pubblicazione: (2025)
di: Wang, Zhican, et al.
Pubblicazione: (2025)
Enthuse: Efficient Adaptable High-throughput Streaming Aggregation Engines
di: Papaphilippou, Philippos, et al.
Pubblicazione: (2024)
di: Papaphilippou, Philippos, et al.
Pubblicazione: (2024)
Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
di: Wu, Haoran, et al.
Pubblicazione: (2025)
di: Wu, Haoran, et al.
Pubblicazione: (2025)
Accelerating Transformer Inference for Translation via Parallel Decoding
di: Santilli, Andrea, et al.
Pubblicazione: (2023)
di: Santilli, Andrea, et al.
Pubblicazione: (2023)
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
di: Wei, Zhepei, et al.
Pubblicazione: (2025)
di: Wei, Zhepei, et al.
Pubblicazione: (2025)
HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
di: Sun, Chang, et al.
Pubblicazione: (2026)
di: Sun, Chang, et al.
Pubblicazione: (2026)
CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
di: Zhou, Enyu, et al.
Pubblicazione: (2025)
di: Zhou, Enyu, et al.
Pubblicazione: (2025)
DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent Method
di: Khaled, Ahmed, et al.
Pubblicazione: (2023)
di: Khaled, Ahmed, et al.
Pubblicazione: (2023)
TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce Edge
di: Kwon, Young D., et al.
Pubblicazione: (2023)
di: Kwon, Young D., et al.
Pubblicazione: (2023)
Vortex‐Beam Multiplexing Emitter Using Advanced Manufactured Leaky Cables
di: Feiyang Deng, et al.
Pubblicazione: (2024)
di: Feiyang Deng, et al.
Pubblicazione: (2024)
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
di: Mishchenko, Konstantin, et al.
Pubblicazione: (2023)
di: Mishchenko, Konstantin, et al.
Pubblicazione: (2023)
Adaptive Proximal Gradient Method for Convex Optimization
di: Malitsky, Yura, et al.
Pubblicazione: (2023)
di: Malitsky, Yura, et al.
Pubblicazione: (2023)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
di: Zhao, Alan, et al.
Pubblicazione: (2026)
di: Zhao, Alan, et al.
Pubblicazione: (2026)
SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
di: Zhang, Hengrui, et al.
Pubblicazione: (2025)
di: Zhang, Hengrui, et al.
Pubblicazione: (2025)
The Future of Consumer Edge-AI Computing
di: Laskaridis, Stefanos, et al.
Pubblicazione: (2022)
di: Laskaridis, Stefanos, et al.
Pubblicazione: (2022)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
di: Xiang, Maoyang, et al.
Pubblicazione: (2025)
di: Xiang, Maoyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Progressive Mixed-Precision Decoding for Efficient LLM Inference
di: Chen, Hao Mark, et al.
Pubblicazione: (2024) -
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
di: Zhang, Zehuan, et al.
Pubblicazione: (2024) -
CARIn: Constraint-Aware and Responsive Inference on Heterogeneous Devices for Single- and Multi-DNN Workloads
di: Panopoulos, Ioannis, et al.
Pubblicazione: (2024) -
Enhancing LLM-based Quantum Code Generation with Multi-Agent Optimization and Quantum Error Correction
di: Campbell, Charlie, et al.
Pubblicazione: (2025) -
MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
di: Nikolaidis, Sokratis, et al.
Pubblicazione: (2024)