High-Rate Quantized Matrix Multiplication II
Fuente:
arXiv
Salvato in:
| Autori principali: | Ordentlich, Or, Polyanskiy, Yury |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimal Quantization for Matrix Multiplication
di: Ordentlich, Or, et al.
Pubblicazione: (2024)
di: Ordentlich, Or, et al.
Pubblicazione: (2024)
High-Rate Quantized Matrix Multiplication I
di: Ordentlich, Or, et al.
Pubblicazione: (2026)
di: Ordentlich, Or, et al.
Pubblicazione: (2026)
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
di: Savkin, Semyon, et al.
Pubblicazione: (2025)
di: Savkin, Semyon, et al.
Pubblicazione: (2025)
Price of universality in vector quantization is at most 0.11 bit
di: Harbuzova, Alina, et al.
Pubblicazione: (2026)
di: Harbuzova, Alina, et al.
Pubblicazione: (2026)
WaterSIC: information-theoretically (near) optimal linear layer quantization
di: Lifar, Egor, et al.
Pubblicazione: (2026)
di: Lifar, Egor, et al.
Pubblicazione: (2026)
High-Rate Nested-Lattice Quantized Matrix Multiplication with Small Lookup Tables
di: Kaplan, Iris, et al.
Pubblicazione: (2025)
di: Kaplan, Iris, et al.
Pubblicazione: (2025)
PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources
di: Park, Bumsu, et al.
Pubblicazione: (2026)
di: Park, Bumsu, et al.
Pubblicazione: (2026)
On the Minimax Regret of Sequential Probability Assignment via Square-Root Entropy
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
Three Quantization Regimes for ReLU Networks
di: Ou, Weigutian, et al.
Pubblicazione: (2024)
di: Ou, Weigutian, et al.
Pubblicazione: (2024)
Optimizing Learned Image Compression on Scalar and Entropy-Constraint Quantization
di: Borzechowski, Florian, et al.
Pubblicazione: (2025)
di: Borzechowski, Florian, et al.
Pubblicazione: (2025)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
di: Lee, Namyoon, et al.
Pubblicazione: (2026)
di: Lee, Namyoon, et al.
Pubblicazione: (2026)
Distributed and Rate-Adaptive Feature Compression
di: Deshmukh, Aditya, et al.
Pubblicazione: (2024)
di: Deshmukh, Aditya, et al.
Pubblicazione: (2024)
Deep Learning and Matrix Completion-aided IoT Network Localization in the Outlier Scenarios
di: Kim, Sunwoo
Pubblicazione: (2025)
di: Kim, Sunwoo
Pubblicazione: (2025)
Semantic Rate Distortion and Posterior Design: Compute Constraints, Multimodality, and Strategic Inference
di: Akyol, Emrah
Pubblicazione: (2026)
di: Akyol, Emrah
Pubblicazione: (2026)
Rotation Invariant Quantization for Model Compression
di: Kampeas, Joseph, et al.
Pubblicazione: (2023)
di: Kampeas, Joseph, et al.
Pubblicazione: (2023)
Statistical Inference with Limited Memory: A Survey
di: Berg, Tomer, et al.
Pubblicazione: (2023)
di: Berg, Tomer, et al.
Pubblicazione: (2023)
Scaling Limits of Long-Context Transformers
di: Bruno, Giuseppe, et al.
Pubblicazione: (2026)
di: Bruno, Giuseppe, et al.
Pubblicazione: (2026)
Optimal Scalar Quantization for Matrix Multiplication: Closed-Form Density and Phase Transition
di: Ang, Calvin, et al.
Pubblicazione: (2026)
di: Ang, Calvin, et al.
Pubblicazione: (2026)
SQuat: Subspace-orthogonal KV Cache Quantization
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
di: Zhao, Qingyue, et al.
Pubblicazione: (2026)
di: Zhao, Qingyue, et al.
Pubblicazione: (2026)
The Voronoi Spherical CDF for Lattices and Linear Codes: New Bounds for Quantization and Coding
di: Ordentlich, Or
Pubblicazione: (2025)
di: Ordentlich, Or
Pubblicazione: (2025)
Representation Alignment Rests on Linear Structure
di: Bangachev, Kiril, et al.
Pubblicazione: (2026)
di: Bangachev, Kiril, et al.
Pubblicazione: (2026)
Turbo-CF: Matrix Decomposition-Free Graph Filtering for Fast Recommendation
di: Park, Jin-Duk, et al.
Pubblicazione: (2024)
di: Park, Jin-Duk, et al.
Pubblicazione: (2024)
Two-Dimensional Quantization for Geometry-Aware Audio Coding
di: Shuster, Tal, et al.
Pubblicazione: (2025)
di: Shuster, Tal, et al.
Pubblicazione: (2025)
ToDMA: Large Model-Driven Token-Domain Multiple Access for Semantic Communications
di: Qiao, Li, et al.
Pubblicazione: (2025)
di: Qiao, Li, et al.
Pubblicazione: (2025)
Modular Learning of Deep Causal Generative Models for High-dimensional Causal Inference
di: Rahman, Md Musfiqur, et al.
Pubblicazione: (2024)
di: Rahman, Md Musfiqur, et al.
Pubblicazione: (2024)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
di: Zimin, Aleksandr, et al.
Pubblicazione: (2026)
di: Zimin, Aleksandr, et al.
Pubblicazione: (2026)
Global Minimizers of Sigmoid Contrastive Loss
di: Bangachev, Kiril, et al.
Pubblicazione: (2025)
di: Bangachev, Kiril, et al.
Pubblicazione: (2025)
Unsourced Multiple Access: A Coding Paradigm for Massive Random Access
di: Liva, Gianluigi, et al.
Pubblicazione: (2024)
di: Liva, Gianluigi, et al.
Pubblicazione: (2024)
Optimal Online Bookmaking for Binary Games
di: Bhatt, Alankrita, et al.
Pubblicazione: (2025)
di: Bhatt, Alankrita, et al.
Pubblicazione: (2025)
On the Limits of Self-Improving in Large Language Models: The Singularity Is Not Near Without Symbolic Model Synthesis
di: Zenil, Hector
Pubblicazione: (2026)
di: Zenil, Hector
Pubblicazione: (2026)
Deep Randomized Distributed Function Computation (DeepRDFC): Neural Distributed Channel Simulation
di: Bergström, Didrik, et al.
Pubblicazione: (2026)
di: Bergström, Didrik, et al.
Pubblicazione: (2026)
Beyond the Loss Curve: Scaling Laws, Active Learning, and the Limits of Learning from Exact Posteriors
di: Khorasani, Arian, et al.
Pubblicazione: (2026)
di: Khorasani, Arian, et al.
Pubblicazione: (2026)
Contextual Control without Memory Growth in a Context-Switching Task
di: Kim, Song-Ju
Pubblicazione: (2026)
di: Kim, Song-Ju
Pubblicazione: (2026)
Why Self-Supervised Encoders Want to Be Normal
di: Domb, Yuval
Pubblicazione: (2026)
di: Domb, Yuval
Pubblicazione: (2026)
The Causal Description Gap: Information-Theoretic Separations Across Pearl's Hierarchy
di: Emadi, Seyed Morteza
Pubblicazione: (2026)
di: Emadi, Seyed Morteza
Pubblicazione: (2026)
A Rational Account of Categorization Based on Information Theory
di: MacLellan, Christopher J., et al.
Pubblicazione: (2026)
di: MacLellan, Christopher J., et al.
Pubblicazione: (2026)
On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training
di: Niu, Xueyan, et al.
Pubblicazione: (2026)
di: Niu, Xueyan, et al.
Pubblicazione: (2026)
The Value of Covariance Matching in Gaussian DDPMs and the Lanczos Sampler
di: Akhtar, Md Sahil, et al.
Pubblicazione: (2026)
di: Akhtar, Md Sahil, et al.
Pubblicazione: (2026)
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
di: Kim, Munsik
Pubblicazione: (2026)
di: Kim, Munsik
Pubblicazione: (2026)
Documenti analoghi
-
Optimal Quantization for Matrix Multiplication
di: Ordentlich, Or, et al.
Pubblicazione: (2024) -
High-Rate Quantized Matrix Multiplication I
di: Ordentlich, Or, et al.
Pubblicazione: (2026) -
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
di: Savkin, Semyon, et al.
Pubblicazione: (2025) -
Price of universality in vector quantization is at most 0.11 bit
di: Harbuzova, Alina, et al.
Pubblicazione: (2026) -
WaterSIC: information-theoretically (near) optimal linear layer quantization
di: Lifar, Egor, et al.
Pubblicazione: (2026)