Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
Fuente:
arXiv
Salvato in:
| Autori principali: | Gautam, Aayush, Gagrani, Mukul, Park, Junyoung, Lee, Mingu, Lott, Chiris, Reddy, Narasimha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
di: Jeon, Wonseok, et al.
Pubblicazione: (2024)
di: Jeon, Wonseok, et al.
Pubblicazione: (2024)
On Speculative Decoding for Multimodal Large Language Models
di: Gagrani, Mukul, et al.
Pubblicazione: (2024)
di: Gagrani, Mukul, et al.
Pubblicazione: (2024)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
di: Goel, Raghavv, et al.
Pubblicazione: (2024)
di: Goel, Raghavv, et al.
Pubblicazione: (2024)
QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill
di: Jones, Dalton, et al.
Pubblicazione: (2026)
di: Jones, Dalton, et al.
Pubblicazione: (2026)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
di: Gautam, Aayush, et al.
Pubblicazione: (2025)
di: Gautam, Aayush, et al.
Pubblicazione: (2025)
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
di: Shrestha, Susav, et al.
Pubblicazione: (2025)
di: Shrestha, Susav, et al.
Pubblicazione: (2025)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
di: Goel, Raghavv, et al.
Pubblicazione: (2025)
di: Goel, Raghavv, et al.
Pubblicazione: (2025)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
di: Agrawal, Sudhanshu, et al.
Pubblicazione: (2025)
di: Agrawal, Sudhanshu, et al.
Pubblicazione: (2025)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
di: Smithline, Gabriel, et al.
Pubblicazione: (2026)
di: Smithline, Gabriel, et al.
Pubblicazione: (2026)
Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing
di: Goel, Raghavv, et al.
Pubblicazione: (2026)
di: Goel, Raghavv, et al.
Pubblicazione: (2026)
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
di: An, Tai, et al.
Pubblicazione: (2025)
di: An, Tai, et al.
Pubblicazione: (2025)
CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
di: Lv, Junlin, et al.
Pubblicazione: (2024)
di: Lv, Junlin, et al.
Pubblicazione: (2024)
Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
di: Song, Chendong, et al.
Pubblicazione: (2026)
di: Song, Chendong, et al.
Pubblicazione: (2026)
KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments
di: Park, Junyoung, et al.
Pubblicazione: (2025)
di: Park, Junyoung, et al.
Pubblicazione: (2025)
UMoE: Unifying Attention and FFN with Shared Experts
di: Yang, Yuanhang, et al.
Pubblicazione: (2025)
di: Yang, Yuanhang, et al.
Pubblicazione: (2025)
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
di: Cai, Yuxuan, et al.
Pubblicazione: (2025)
di: Cai, Yuxuan, et al.
Pubblicazione: (2025)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
di: Goel, Raghavv, et al.
Pubblicazione: (2025)
di: Goel, Raghavv, et al.
Pubblicazione: (2025)
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
di: Lee, Kwanhee, et al.
Pubblicazione: (2025)
di: Lee, Kwanhee, et al.
Pubblicazione: (2025)
LLM Serving Optimization with Variable Prefill and Decode Lengths
di: Wang, Meixuan, et al.
Pubblicazione: (2025)
di: Wang, Meixuan, et al.
Pubblicazione: (2025)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
di: Pei, Zehua, et al.
Pubblicazione: (2025)
di: Pei, Zehua, et al.
Pubblicazione: (2025)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
di: Peng, Dan, et al.
Pubblicazione: (2025)
di: Peng, Dan, et al.
Pubblicazione: (2025)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
di: Ji, Xiaodong, et al.
Pubblicazione: (2025)
di: Ji, Xiaodong, et al.
Pubblicazione: (2025)
ConFu: Contemplate the Future for Better Speculative Sampling
di: Qin, Zongyue, et al.
Pubblicazione: (2026)
di: Qin, Zongyue, et al.
Pubblicazione: (2026)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
di: Qiao, Aurick, et al.
Pubblicazione: (2024)
di: Qiao, Aurick, et al.
Pubblicazione: (2024)
Steered LLM Activations are Non-Surjective
di: Mishra, Aayush, et al.
Pubblicazione: (2026)
di: Mishra, Aayush, et al.
Pubblicazione: (2026)
MaD-Scientist: AI-based Scientist solving Convection-Diffusion-Reaction Equations Using Massive PINN-Based Prior Data
di: Kang, Mingu, et al.
Pubblicazione: (2024)
di: Kang, Mingu, et al.
Pubblicazione: (2024)
Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models
di: Zhao, Siyan, et al.
Pubblicazione: (2024)
di: Zhao, Siyan, et al.
Pubblicazione: (2024)
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting
di: Zhao, Yanjun, et al.
Pubblicazione: (2024)
di: Zhao, Yanjun, et al.
Pubblicazione: (2024)
Block-Attention for Efficient Prefilling
di: Ma, Dongyang, et al.
Pubblicazione: (2024)
di: Ma, Dongyang, et al.
Pubblicazione: (2024)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
di: Liu, Ningyuan, et al.
Pubblicazione: (2025)
di: Liu, Ningyuan, et al.
Pubblicazione: (2025)
Sparsity and Out-of-Distribution Generalization
di: Aaronson, Scott, et al.
Pubblicazione: (2026)
di: Aaronson, Scott, et al.
Pubblicazione: (2026)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
FwdLLM: Efficient FedLLM using Forward Gradient
di: Xu, Mengwei, et al.
Pubblicazione: (2023)
di: Xu, Mengwei, et al.
Pubblicazione: (2023)
Can We Predict the Unpredictable? Leveraging DisasterNet-LLM for Multimodal Disaster Classification
di: Kulahara, Manaswi, et al.
Pubblicazione: (2025)
di: Kulahara, Manaswi, et al.
Pubblicazione: (2025)
DECODE: Data-driven Energy Consumption Prediction leveraging Historical Data and Environmental Factors in Buildings
di: Mishra, Aditya, et al.
Pubblicazione: (2023)
di: Mishra, Aditya, et al.
Pubblicazione: (2023)
VSPrefill: Vertical-Slash Sparse Attention with Lightweight Indexing for Long-Context Prefilling
di: Guanzhong, Chen
Pubblicazione: (2026)
di: Guanzhong, Chen
Pubblicazione: (2026)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
di: Tang, Xiaojuan, et al.
Pubblicazione: (2025)
di: Tang, Xiaojuan, et al.
Pubblicazione: (2025)
Spark Transformer: Reactivating Sparsity in FFN and Attention
di: You, Chong, et al.
Pubblicazione: (2025)
di: You, Chong, et al.
Pubblicazione: (2025)
Shaping Zero-Shot Coordination via State Blocking
di: Kang, Mingu, et al.
Pubblicazione: (2026)
di: Kang, Mingu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
di: Jeon, Wonseok, et al.
Pubblicazione: (2024) -
On Speculative Decoding for Multimodal Large Language Models
di: Gagrani, Mukul, et al.
Pubblicazione: (2024) -
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
di: Goel, Raghavv, et al.
Pubblicazione: (2024) -
QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill
di: Jones, Dalton, et al.
Pubblicazione: (2026) -
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
di: Gautam, Aayush, et al.
Pubblicazione: (2025)