SVD Contextual Sparsity Predictors for Fast LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Serbin, Georgii, Koshkin, Kirill, Sun, Zhongao, Bistrigova, Anastasiya, Korikov, C. C. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
by: Shrestha, Susav, et al.
Published: (2025)
by: Shrestha, Susav, et al.
Published: (2025)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
DipSVD: Dual-importance Protected SVD for Efficient LLM Compression
by: Ding, Xuan, et al.
Published: (2025)
by: Ding, Xuan, et al.
Published: (2025)
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
by: Khaki, Samir, et al.
Published: (2025)
by: Khaki, Samir, et al.
Published: (2025)
Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference
by: Joo, Donghyeon, et al.
Published: (2025)
by: Joo, Donghyeon, et al.
Published: (2025)
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
by: Gautam, Aayush, et al.
Published: (2026)
by: Gautam, Aayush, et al.
Published: (2026)
High-dimensional Contextual Bandit Problem without Sparsity
by: Komiyama, Junpei, et al.
Published: (2023)
by: Komiyama, Junpei, et al.
Published: (2023)
Activity Sparsity Complements Weight Sparsity for Efficient RNN Inference
by: Mukherji, Rishav, et al.
Published: (2023)
by: Mukherji, Rishav, et al.
Published: (2023)
Robustness questions the interpretability of graph neural networks: what to do?
by: Lukyanov, Kirill, et al.
Published: (2025)
by: Lukyanov, Kirill, et al.
Published: (2025)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)
by: Tang, Jiaming, et al.
Published: (2024)
SVD-NO: Learning PDE Solution Operators with SVD Integral Kernels
by: Koren, Noam, et al.
Published: (2025)
by: Koren, Noam, et al.
Published: (2025)
Navigating Sparsities in High-Dimensional Linear Contextual Bandits
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression
by: Sinha, Atul Kumar, et al.
Published: (2026)
by: Sinha, Atul Kumar, et al.
Published: (2026)
CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models
by: Lee, Donghyun, et al.
Published: (2024)
by: Lee, Donghyun, et al.
Published: (2024)
skscope: Fast Sparsity-Constrained Optimization in Python
by: Wang, Zezhi, et al.
Published: (2024)
by: Wang, Zezhi, et al.
Published: (2024)
Framework GNN-AID: Graph Neural Network Analysis Interpretation and Defense
by: Lukyanov, Kirill, et al.
Published: (2025)
by: Lukyanov, Kirill, et al.
Published: (2025)
WiSparse: Boosting LLM Inference Efficiency with Weight-Aware Mixed Activation Sparsity
by: Chen, Lei, et al.
Published: (2026)
by: Chen, Lei, et al.
Published: (2026)
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
MOSAIC: Minimax-Optimal Sparsity-Adaptive Inference for Change Points in Dynamic Networks
by: Fan, Yingying, et al.
Published: (2025)
by: Fan, Yingying, et al.
Published: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
Beyond Conformal Predictors: Adaptive Conformal Inference with Confidence Predictors
by: Szabadváry, Johan Hallberg, et al.
Published: (2024)
by: Szabadváry, Johan Hallberg, et al.
Published: (2024)
Beyond Uniform SVD:Dual-Level Optimization across Columns and Modules for LLM Compression
by: Xv, Lin, et al.
Published: (2025)
by: Xv, Lin, et al.
Published: (2025)
Inference Time Context Sparsity: Illusion or Opportunity?
by: Joshi, Sahil, et al.
Published: (2026)
by: Joshi, Sahil, et al.
Published: (2026)
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models
by: Hong, Chengjie, et al.
Published: (2026)
by: Hong, Chengjie, et al.
Published: (2026)
PCA, SVD, and Centering of Data
by: Kim, Donggun, et al.
Published: (2023)
by: Kim, Donggun, et al.
Published: (2023)
LLM Sparsity Prior for Robust Feature Selection
by: Skinner, Caleb, et al.
Published: (2026)
by: Skinner, Caleb, et al.
Published: (2026)
GQSA: Group Quantization and Sparsity for Accelerating Large Language Model Inference
by: Zeng, Chao, et al.
Published: (2024)
by: Zeng, Chao, et al.
Published: (2024)
Federated Learning via Variational Bayesian Inference: Personalization, Sparsity and Clustering
by: Zhang, Xu, et al.
Published: (2023)
by: Zhang, Xu, et al.
Published: (2023)
Inverted Activations: Reducing Memory Footprint in Neural Network Training
by: Novikov, Georgii, et al.
Published: (2024)
by: Novikov, Georgii, et al.
Published: (2024)
Training Without Orthogonalization, Inference With SVD: A Gradient Analysis of Rotation Representations
by: Choy, Chris
Published: (2026)
by: Choy, Chris
Published: (2026)
Fast Best-in-Class Regret for Contextual Bandits
by: Girard, Samuel, et al.
Published: (2025)
by: Girard, Samuel, et al.
Published: (2025)
KLLM: Fast LLM Inference with K-Means Quantization
by: Wu, Xueying, et al.
Published: (2025)
by: Wu, Xueying, et al.
Published: (2025)
Different Prompts, Different Ranks: Prompt-aware Dynamic Rank Selection for SVD-based LLM Compression
by: Zhu, Hengyi, et al.
Published: (2026)
by: Zhu, Hengyi, et al.
Published: (2026)
$k$-SVD with Gradient Descent
by: Jedra, Yassir, et al.
Published: (2025)
by: Jedra, Yassir, et al.
Published: (2025)
Deep Neural Network Initialization with Sparsity Inducing Activations
by: Price, Ilan, et al.
Published: (2024)
by: Price, Ilan, et al.
Published: (2024)
Similar Items
-
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
by: Shrestha, Susav, et al.
Published: (2025) -
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
by: Akhauri, Yash, et al.
Published: (2024) -
DipSVD: Dual-importance Protected SVD for Efficient LLM Compression
by: Ding, Xuan, et al.
Published: (2025) -
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
by: Khaki, Samir, et al.
Published: (2025) -
Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
by: Wang, Qinsi, et al.
Published: (2025)