NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Savkin, Semyon, Porat, Eitan, Ordentlich, Or, Polyanskiy, Yury |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
High-Rate Quantized Matrix Multiplication II
por: Ordentlich, Or, et al.
Publicado: (2026)
por: Ordentlich, Or, et al.
Publicado: (2026)
Optimal Quantization for Matrix Multiplication
por: Ordentlich, Or, et al.
Publicado: (2024)
por: Ordentlich, Or, et al.
Publicado: (2024)
WaterSIC: information-theoretically (near) optimal linear layer quantization
por: Lifar, Egor, et al.
Publicado: (2026)
por: Lifar, Egor, et al.
Publicado: (2026)
High-Rate Quantized Matrix Multiplication I
por: Ordentlich, Or, et al.
Publicado: (2026)
por: Ordentlich, Or, et al.
Publicado: (2026)
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
por: Xie, Jianhang, et al.
Publicado: (2025)
por: Xie, Jianhang, et al.
Publicado: (2025)
High-Rate Nested-Lattice Quantized Matrix Multiplication with Small Lookup Tables
por: Kaplan, Iris, et al.
Publicado: (2025)
por: Kaplan, Iris, et al.
Publicado: (2025)
Price of universality in vector quantization is at most 0.11 bit
por: Harbuzova, Alina, et al.
Publicado: (2026)
por: Harbuzova, Alina, et al.
Publicado: (2026)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
por: Lee, Namyoon, et al.
Publicado: (2026)
por: Lee, Namyoon, et al.
Publicado: (2026)
PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources
por: Park, Bumsu, et al.
Publicado: (2026)
por: Park, Bumsu, et al.
Publicado: (2026)
Generalized Nested Latent Variable Models for Lossy Coding applied to Wind Turbine Scenarios
por: Pérez-Gonzalo, Raül, et al.
Publicado: (2024)
por: Pérez-Gonzalo, Raül, et al.
Publicado: (2024)
The Voronoi Spherical CDF for Lattices and Linear Codes: New Bounds for Quantization and Coding
por: Ordentlich, Or
Publicado: (2025)
por: Ordentlich, Or
Publicado: (2025)
On the Minimax Regret of Sequential Probability Assignment via Square-Root Entropy
por: Jia, Zeyu, et al.
Publicado: (2025)
por: Jia, Zeyu, et al.
Publicado: (2025)
Three Quantization Regimes for ReLU Networks
por: Ou, Weigutian, et al.
Publicado: (2024)
por: Ou, Weigutian, et al.
Publicado: (2024)
Nested Construction of Polar Codes via Transformers
por: Ankireddy, Sravan Kumar, et al.
Publicado: (2024)
por: Ankireddy, Sravan Kumar, et al.
Publicado: (2024)
Optimizing Learned Image Compression on Scalar and Entropy-Constraint Quantization
por: Borzechowski, Florian, et al.
Publicado: (2025)
por: Borzechowski, Florian, et al.
Publicado: (2025)
Compute-Update Federated Learning: A Lattice Coding Approach Over-the-Air
por: Azimi-Abarghouyi, Seyed Mohammad, et al.
Publicado: (2024)
por: Azimi-Abarghouyi, Seyed Mohammad, et al.
Publicado: (2024)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
Deep Learning and Matrix Completion-aided IoT Network Localization in the Outlier Scenarios
por: Kim, Sunwoo
Publicado: (2025)
por: Kim, Sunwoo
Publicado: (2025)
Modeling Churn in Recommender Systems with Aggregated Preferences
por: Keinan, Gur, et al.
Publicado: (2025)
por: Keinan, Gur, et al.
Publicado: (2025)
Rotation Invariant Quantization for Model Compression
por: Kampeas, Joseph, et al.
Publicado: (2023)
por: Kampeas, Joseph, et al.
Publicado: (2023)
DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer Arithmetic
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2025)
por: Shalby, Hazem Hesham Yousef, et al.
Publicado: (2025)
Statistical Inference with Limited Memory: A Survey
por: Berg, Tomer, et al.
Publicado: (2023)
por: Berg, Tomer, et al.
Publicado: (2023)
Scaling Limits of Long-Context Transformers
por: Bruno, Giuseppe, et al.
Publicado: (2026)
por: Bruno, Giuseppe, et al.
Publicado: (2026)
Tele-LLMs: A Series of Specialized Large Language Models for Telecommunications
por: Maatouk, Ali, et al.
Publicado: (2024)
por: Maatouk, Ali, et al.
Publicado: (2024)
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
por: Ouyang, Xu, et al.
Publicado: (2026)
por: Ouyang, Xu, et al.
Publicado: (2026)
Haiku to Opus in Just 10 bits: LLMs Unlock Massive Compression Gains
por: Rinberg, Roy, et al.
Publicado: (2026)
por: Rinberg, Roy, et al.
Publicado: (2026)
SQuat: Subspace-orthogonal KV Cache Quantization
por: Wang, Hao, et al.
Publicado: (2025)
por: Wang, Hao, et al.
Publicado: (2025)
ProductAE: Toward Deep Learning Driven Error-Correction Codes of Large Dimensions
por: Jamali, Mohammad Vahid, et al.
Publicado: (2023)
por: Jamali, Mohammad Vahid, et al.
Publicado: (2023)
Representation Alignment Rests on Linear Structure
por: Bangachev, Kiril, et al.
Publicado: (2026)
por: Bangachev, Kiril, et al.
Publicado: (2026)
The Radio-Frequency Transformer for Signal Separation
por: Lifar, Egor, et al.
Publicado: (2026)
por: Lifar, Egor, et al.
Publicado: (2026)
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
por: Song, Jaewoo, et al.
Publicado: (2025)
por: Song, Jaewoo, et al.
Publicado: (2025)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
por: Zhao, Zhixiong, et al.
Publicado: (2025)
por: Zhao, Zhixiong, et al.
Publicado: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
por: Yan, Xianglong, et al.
Publicado: (2026)
por: Yan, Xianglong, et al.
Publicado: (2026)
Machine Learning Models to Identify Promising Nested Antiresonance Nodeless Fiber Designs
por: Eltaieb, Rania A., et al.
Publicado: (2026)
por: Eltaieb, Rania A., et al.
Publicado: (2026)
Turbo-CF: Matrix Decomposition-Free Graph Filtering for Fast Recommendation
por: Park, Jin-Duk, et al.
Publicado: (2024)
por: Park, Jin-Duk, et al.
Publicado: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
por: Shen, Xuan, et al.
Publicado: (2023)
por: Shen, Xuan, et al.
Publicado: (2023)
Two-Dimensional Quantization for Geometry-Aware Audio Coding
por: Shuster, Tal, et al.
Publicado: (2025)
por: Shuster, Tal, et al.
Publicado: (2025)
PolarQuant: Quantizing KV Caches with Polar Transformation
por: Han, Insu, et al.
Publicado: (2025)
por: Han, Insu, et al.
Publicado: (2025)
SPEX: Scaling Feature Interaction Explanations for LLMs
por: Kang, Justin Singh, et al.
Publicado: (2025)
por: Kang, Justin Singh, et al.
Publicado: (2025)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
por: Zhang, Jinghe, et al.
Publicado: (2026)
por: Zhang, Jinghe, et al.
Publicado: (2026)
Ejemplares similares
-
High-Rate Quantized Matrix Multiplication II
por: Ordentlich, Or, et al.
Publicado: (2026) -
Optimal Quantization for Matrix Multiplication
por: Ordentlich, Or, et al.
Publicado: (2024) -
WaterSIC: information-theoretically (near) optimal linear layer quantization
por: Lifar, Egor, et al.
Publicado: (2026) -
High-Rate Quantized Matrix Multiplication I
por: Ordentlich, Or, et al.
Publicado: (2026) -
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
por: Xie, Jianhang, et al.
Publicado: (2025)