Progressive Mixed-Precision Decoding for Efficient LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Hao Mark, Tan, Fuwen, Kouris, Alexandros, Lee, Royson, Fan, Hongxiang, Venieris, Stylianos I. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
Speculative Decoding with a Speculative Vocabulary
von: Williams, Miles, et al.
Veröffentlicht: (2026)
von: Williams, Miles, et al.
Veröffentlicht: (2026)
The Future of Consumer Edge-AI Computing
von: Laskaridis, Stefanos, et al.
Veröffentlicht: (2022)
von: Laskaridis, Stefanos, et al.
Veröffentlicht: (2022)
CARIn: Constraint-Aware and Responsive Inference on Heterogeneous Devices for Single- and Multi-DNN Workloads
von: Panopoulos, Ioannis, et al.
Veröffentlicht: (2024)
von: Panopoulos, Ioannis, et al.
Veröffentlicht: (2024)
MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
von: Nikolaidis, Sokratis, et al.
Veröffentlicht: (2024)
von: Nikolaidis, Sokratis, et al.
Veröffentlicht: (2024)
FedP$^2$EFT: Federated Learning to Personalize PEFT for Multilingual LLMs
von: Lee, Royson, et al.
Veröffentlicht: (2025)
von: Lee, Royson, et al.
Veröffentlicht: (2025)
Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
von: Fu, Yichao, et al.
Veröffentlicht: (2024)
von: Fu, Yichao, et al.
Veröffentlicht: (2024)
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
von: Cai, Tianle, et al.
Veröffentlicht: (2024)
von: Cai, Tianle, et al.
Veröffentlicht: (2024)
Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generation
von: Zhang, Hongxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Hongxiang, et al.
Veröffentlicht: (2025)
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
von: Fan, Zehao, et al.
Veröffentlicht: (2025)
von: Fan, Zehao, et al.
Veröffentlicht: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
von: Li, Xing, et al.
Veröffentlicht: (2025)
von: Li, Xing, et al.
Veröffentlicht: (2025)
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
Universal Model Routing for Efficient LLM Inference
von: Jitkrittum, Wittawat, et al.
Veröffentlicht: (2025)
von: Jitkrittum, Wittawat, et al.
Veröffentlicht: (2025)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
von: Ma, Xuezhe, et al.
Veröffentlicht: (2024)
von: Ma, Xuezhe, et al.
Veröffentlicht: (2024)
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference
von: Jang, Wonsuk, et al.
Veröffentlicht: (2025)
von: Jang, Wonsuk, et al.
Veröffentlicht: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
von: Qin, Zongyue, et al.
Veröffentlicht: (2024)
von: Qin, Zongyue, et al.
Veröffentlicht: (2024)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
von: Yang, Seongjun, et al.
Veröffentlicht: (2023)
von: Yang, Seongjun, et al.
Veröffentlicht: (2023)
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
von: Ren, Liliang, et al.
Veröffentlicht: (2025)
von: Ren, Liliang, et al.
Veröffentlicht: (2025)
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Enhancing Trustworthiness with Mixed Precision: Benchmarks, Opportunities, and Challenges
von: Lu, Guanxi, et al.
Veröffentlicht: (2025)
von: Lu, Guanxi, et al.
Veröffentlicht: (2025)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
SplaXBERT: Leveraging Mixed Precision Training and Context Splitting for Question Answering
von: Yufan, Zhu, et al.
Veröffentlicht: (2024)
von: Yufan, Zhu, et al.
Veröffentlicht: (2024)
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
von: Hong, Yinrong, et al.
Veröffentlicht: (2025)
von: Hong, Yinrong, et al.
Veröffentlicht: (2025)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
von: Wu, Zimeng, et al.
Veröffentlicht: (2026)
von: Wu, Zimeng, et al.
Veröffentlicht: (2026)
PGF-Net: A Progressive Gated-Fusion Framework for Efficient Multimodal Sentiment Analysis
von: Wen, Bin, et al.
Veröffentlicht: (2025)
von: Wen, Bin, et al.
Veröffentlicht: (2025)
Entropy Adaptive Decoding: Dynamic Model Switching for Efficient Inference
von: Simonds, Toby
Veröffentlicht: (2025)
von: Simonds, Toby
Veröffentlicht: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
von: Zhang, Zeliang, et al.
Veröffentlicht: (2026)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2026)
FlashDecoding++: Faster Large Language Model Inference on GPUs
von: Hong, Ke, et al.
Veröffentlicht: (2023)
von: Hong, Ke, et al.
Veröffentlicht: (2023)
Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
von: Shen, Ethan, et al.
Veröffentlicht: (2024)
von: Shen, Ethan, et al.
Veröffentlicht: (2024)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
von: Kim, Jang-Hyun, et al.
Veröffentlicht: (2026)
von: Kim, Jang-Hyun, et al.
Veröffentlicht: (2026)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
RelayLLM: Efficient Reasoning via Collaborative Decoding
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
von: Huang, Chengsong, et al.
Veröffentlicht: (2026)
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
LIDS: LLM Summary Inference Under the Layered Lens
von: Park, Dylan, et al.
Veröffentlicht: (2026)
von: Park, Dylan, et al.
Veröffentlicht: (2026)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
von: Deng, Jianing, et al.
Veröffentlicht: (2026)
von: Deng, Jianing, et al.
Veröffentlicht: (2026)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
von: Federici, Marco, et al.
Veröffentlicht: (2024)
von: Federici, Marco, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024) -
Speculative Decoding with a Speculative Vocabulary
von: Williams, Miles, et al.
Veröffentlicht: (2026) -
The Future of Consumer Edge-AI Computing
von: Laskaridis, Stefanos, et al.
Veröffentlicht: (2022) -
CARIn: Constraint-Aware and Responsive Inference on Heterogeneous Devices for Single- and Multi-DNN Workloads
von: Panopoulos, Ioannis, et al.
Veröffentlicht: (2024) -
MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
von: Nikolaidis, Sokratis, et al.
Veröffentlicht: (2024)