SDQ: Sparse Decomposed Quantization for LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Jeong, Geonhwa, Tsai, Po-An, Keckler, Stephen W., Krishna, Tushar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024)
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
di: Kang, Hao, et al.
Pubblicazione: (2024)
di: Kang, Hao, et al.
Pubblicazione: (2024)
Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
di: Xia, Junhao, et al.
Pubblicazione: (2025)
di: Xia, Junhao, et al.
Pubblicazione: (2025)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
Breaking the Blocks: Continuous Low-Rank Decomposed Scaling for Unified LLM Quantization and Adaptation
di: Tang, Pingzhi, et al.
Pubblicazione: (2026)
di: Tang, Pingzhi, et al.
Pubblicazione: (2026)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
di: Chhugani, Jatin, et al.
Pubblicazione: (2026)
di: Chhugani, Jatin, et al.
Pubblicazione: (2026)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
di: Haziza, Daniel, et al.
Pubblicazione: (2025)
di: Haziza, Daniel, et al.
Pubblicazione: (2025)
Why Inference in Large Models Becomes Decomposable After Training
di: Jin, Jidong
Pubblicazione: (2026)
di: Jin, Jidong
Pubblicazione: (2026)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
di: Gautam, Arpit Singh, et al.
Pubblicazione: (2026)
di: Gautam, Arpit Singh, et al.
Pubblicazione: (2026)
4bit-Quantization in Vector-Embedding for RAG
di: Jeong, Taehee
Pubblicazione: (2025)
di: Jeong, Taehee
Pubblicazione: (2025)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
di: Zhou, Ruijie, et al.
Pubblicazione: (2026)
di: Zhou, Ruijie, et al.
Pubblicazione: (2026)
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
di: Le, Hoang Anh Duy, et al.
Pubblicazione: (2026)
di: Le, Hoang Anh Duy, et al.
Pubblicazione: (2026)
WiSparse: Boosting LLM Inference Efficiency with Weight-Aware Mixed Activation Sparsity
di: Chen, Lei, et al.
Pubblicazione: (2026)
di: Chen, Lei, et al.
Pubblicazione: (2026)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
di: Zhang, Hailin, et al.
Pubblicazione: (2024)
di: Zhang, Hailin, et al.
Pubblicazione: (2024)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
di: Zhang, Yu, et al.
Pubblicazione: (2024)
di: Zhang, Yu, et al.
Pubblicazione: (2024)
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
di: Hajimolahoseini, Habib, et al.
Pubblicazione: (2023)
di: Hajimolahoseini, Habib, et al.
Pubblicazione: (2023)
Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees
di: Chowdhury, Mohammed Nowaz Rabbani, et al.
Pubblicazione: (2026)
di: Chowdhury, Mohammed Nowaz Rabbani, et al.
Pubblicazione: (2026)
On the Sustainability of AI Inferences in the Edge
di: Sobhani, Ghazal, et al.
Pubblicazione: (2025)
di: Sobhani, Ghazal, et al.
Pubblicazione: (2025)
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
di: Wang, Ziyi, et al.
Pubblicazione: (2025)
XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference
di: Witt, Thomas
Pubblicazione: (2026)
di: Witt, Thomas
Pubblicazione: (2026)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
di: Becking, Daniel, et al.
Pubblicazione: (2021)
di: Becking, Daniel, et al.
Pubblicazione: (2021)
Adaptive Weighted Loss for Sequential Recommendations on Sparse Domains
di: Mittal, Akshay, et al.
Pubblicazione: (2025)
di: Mittal, Akshay, et al.
Pubblicazione: (2025)
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
di: Kundurthy, Srivatsa, et al.
Pubblicazione: (2026)
di: Kundurthy, Srivatsa, et al.
Pubblicazione: (2026)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
di: Tsujimura, Hikaru, et al.
Pubblicazione: (2025)
di: Tsujimura, Hikaru, et al.
Pubblicazione: (2025)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
di: Tan, Qitao, et al.
Pubblicazione: (2025)
di: Tan, Qitao, et al.
Pubblicazione: (2025)
Exploiting LLM Quantization
di: Egashira, Kazuki, et al.
Pubblicazione: (2024)
di: Egashira, Kazuki, et al.
Pubblicazione: (2024)
Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor
di: Li, Xiaocan, et al.
Pubblicazione: (2026)
di: Li, Xiaocan, et al.
Pubblicazione: (2026)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
di: Chu, Kexin, et al.
Pubblicazione: (2025)
di: Chu, Kexin, et al.
Pubblicazione: (2025)
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting
di: Zhao, Yanjun, et al.
Pubblicazione: (2024)
di: Zhao, Yanjun, et al.
Pubblicazione: (2024)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
di: Hu, Xing, et al.
Pubblicazione: (2024)
di: Hu, Xing, et al.
Pubblicazione: (2024)
MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents
di: Levinson, Matthew
Pubblicazione: (2026)
di: Levinson, Matthew
Pubblicazione: (2026)
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
di: Zhan, Xiaohua, et al.
Pubblicazione: (2026)
di: Zhan, Xiaohua, et al.
Pubblicazione: (2026)
ICQuant: Index Coding enables Low-bit LLM Quantization
di: Li, Xinlin, et al.
Pubblicazione: (2025)
di: Li, Xinlin, et al.
Pubblicazione: (2025)
Decomposing and Editing Predictions by Modeling Model Computation
di: Shah, Harshay, et al.
Pubblicazione: (2024)
di: Shah, Harshay, et al.
Pubblicazione: (2024)
Decomposing Epistemic Uncertainty for Causal Decision Making
di: Rahman, Md Musfiqur, et al.
Pubblicazione: (2026)
di: Rahman, Md Musfiqur, et al.
Pubblicazione: (2026)
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
di: Hamilton, Sil, et al.
Pubblicazione: (2025)
di: Hamilton, Sil, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024) -
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
di: Kang, Hao, et al.
Pubblicazione: (2024) -
Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024) -
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024) -
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)