Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Lingchao, Fan, Yuwei, Li, Jun, Hu, Chengqiu, Liao, Qichen, Fan, Junyi, Shi, Rui, Miao, Fangzheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025)
by: Liao, Qichen, et al.
Published: (2025)
TP-Aware Dequantization
by: Hoque, Adnan, et al.
Published: (2024)
by: Hoque, Adnan, et al.
Published: (2024)
Dequantization Barriers for Guided Stoquastic Hamiltonians
by: Hamoudi, Yassine, et al.
Published: (2026)
by: Hamoudi, Yassine, et al.
Published: (2026)
Dequantization and Hardness of Spectral Sum Estimation
by: Edenhofer, Roman, et al.
Published: (2025)
by: Edenhofer, Roman, et al.
Published: (2025)
AXELRAM: Quantize Once, Never Dequantize
by: Nishida, Yasushi
Published: (2026)
by: Nishida, Yasushi
Published: (2026)
Dequantization and Color Transfer with Diffusion Models
by: Vavilala, Vaibhav, et al.
Published: (2023)
by: Vavilala, Vaibhav, et al.
Published: (2023)
Dequantizing Short-Path Quantum Algorithms
by: Gall, François Le, et al.
Published: (2026)
by: Gall, François Le, et al.
Published: (2026)
Fast NF4 Dequantization Kernels for Large Language Model Inference
by: Qi, Xiangbo, et al.
Published: (2026)
by: Qi, Xiangbo, et al.
Published: (2026)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
by: Li, Jinhao, et al.
Published: (2023)
by: Li, Jinhao, et al.
Published: (2023)
Improving Detail in Pluralistic Image Inpainting with Feature Dequantization
by: Park, Kyungri, et al.
Published: (2024)
by: Park, Kyungri, et al.
Published: (2024)
On the Quantization-Dequantization Correspondence for (co)Poisson Hopf Algebras
by: Rivezzi, Andrea, et al.
Published: (2026)
by: Rivezzi, Andrea, et al.
Published: (2026)
A Dequantized Algorithm for the Guided Local Hamiltonian Problem
by: Zhang, Yukun, et al.
Published: (2024)
by: Zhang, Yukun, et al.
Published: (2024)
On Dequantization of Supervised Quantum Machine Learning via Random Fourier Features
by: Sahebi, Mehrad, et al.
Published: (2025)
by: Sahebi, Mehrad, et al.
Published: (2025)
New Limits on Distributed Quantum Advantage: Dequantizing Linear Programs
by: Balliu, Alkida, et al.
Published: (2025)
by: Balliu, Alkida, et al.
Published: (2025)
Dequantizing quantum machine learning models using tensor networks
by: Shin, Seongwook, et al.
Published: (2023)
by: Shin, Seongwook, et al.
Published: (2023)
Dequantization of a signal from two parallel quantized observations
by: Kovanda, Vojtěch, et al.
Published: (2024)
by: Kovanda, Vojtěch, et al.
Published: (2024)
DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer Arithmetic
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
Tensor Network Formulation of Dequantized Algorithms for Ground State Energy Estimation
by: Manabe, Hidetaka, et al.
Published: (2025)
by: Manabe, Hidetaka, et al.
Published: (2025)
Robust Dequantization of the Quantum Singular value Transformation and Quantum Machine Learning Algorithms
by: Gall, François Le
Published: (2023)
by: Gall, François Le
Published: (2023)
LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model
by: Shao, Wei, et al.
Published: (2025)
by: Shao, Wei, et al.
Published: (2025)
On the quantum computational complexity of classical linear dynamics with geometrically local interactions: Dequantization and universality
by: Sakamoto, Kazuki, et al.
Published: (2025)
by: Sakamoto, Kazuki, et al.
Published: (2025)
Quantum Inspiration, Classical Advantage: Dequantized particle algorithm for the nonlinear Vlasov-Poisson system
by: Qin, Hong, et al.
Published: (2025)
by: Qin, Hong, et al.
Published: (2025)
Energy-Efficient and Dequantization-Free Q-LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
Dequantizing the Quantum Singular Value Transformation: Hardness and Applications to Quantum Chemistry and the Quantum PCP Conjecture
by: Gharibian, Sevag, et al.
Published: (2021)
by: Gharibian, Sevag, et al.
Published: (2021)
SteeringDiffusion: A Bottlenecked Activation Control Interface for Diffusion Models
by: Wu, Fangzheng, et al.
Published: (2026)
by: Wu, Fangzheng, et al.
Published: (2026)
Online Pseudo-average Shifting Attention(PASA) for Robust Low-precision LLM Inference: Algorithms and Numerical Analysis
by: Cheng, Long, et al.
Published: (2025)
by: Cheng, Long, et al.
Published: (2025)
AIS: Adaptive Importance Sampling for Quantized RL
by: Zhou, Jiajun, et al.
Published: (2026)
by: Zhou, Jiajun, et al.
Published: (2026)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
by: Bu, Tianci, et al.
Published: (2026)
by: Bu, Tianci, et al.
Published: (2026)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
by: Ning, Rui, et al.
Published: (2026)
by: Ning, Rui, et al.
Published: (2026)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
by: Zhao, Alan, et al.
Published: (2026)
by: Zhao, Alan, et al.
Published: (2026)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024)
by: Fu, Qichen, et al.
Published: (2024)
QUAD: Quantization and Parameter-Efficient Tuning of LLM with Activation Decomposition
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
by: Meng, William, et al.
Published: (2025)
by: Meng, William, et al.
Published: (2025)
Bias-Corrected Joint Spectral Embedding for Multilayer Networks with Invariant Subspace: Entrywise Eigenvector Perturbation and Inference
by: Xie, Fangzheng
Published: (2024)
by: Xie, Fangzheng
Published: (2024)
iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM Inference
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
WaferLLM: Large Language Model Inference at Wafer Scale
by: He, Congjie, et al.
Published: (2025)
by: He, Congjie, et al.
Published: (2025)
Flexible Concept Bottleneck Model
by: Du, Xingbo, et al.
Published: (2025)
by: Du, Xingbo, et al.
Published: (2025)
DISC: Dynamic Decomposition Improves LLM Inference Scaling
by: Light, Jonathan, et al.
Published: (2025)
by: Light, Jonathan, et al.
Published: (2025)
CD1a affects the recurrence and prognosis of ovarian cancer
by: Qiong Zhu, et al.
Published: (2024)
by: Qiong Zhu, et al.
Published: (2024)
Similar Items
-
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025) -
TP-Aware Dequantization
by: Hoque, Adnan, et al.
Published: (2024) -
Dequantization Barriers for Guided Stoquastic Hamiltonians
by: Hamoudi, Yassine, et al.
Published: (2026) -
Dequantization and Hardness of Spectral Sum Estimation
by: Edenhofer, Roman, et al.
Published: (2025) -
AXELRAM: Quantize Once, Never Dequantize
by: Nishida, Yasushi
Published: (2026)