Saved in:
| Main Authors: | Müller, Lorenz K., Bich, Philippe, Zhuang, Jiawei, Çelik, Ahmet, Benfenati, Luca, Cavigelli, Lukas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.22944 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
SSSD: Simply-Scalable Speculative Decoding
by: Marzollo, Michele, et al.
Published: (2024)
by: Marzollo, Michele, et al.
Published: (2024)
Don't be so Stief! Learning KV Cache low-rank approximation over the Stiefel manifold
by: Benfenati, Luca, et al.
Published: (2026)
by: Benfenati, Luca, et al.
Published: (2026)
Boosting keyword spotting through on-device learnable user speech characteristics
by: Cioflan, Cristian, et al.
Published: (2024)
by: Cioflan, Cristian, et al.
Published: (2024)
On-Device Domain Learning for Keyword Spotting on Low-Power Extreme Edge Embedded Systems
by: Cioflan, Cristian, et al.
Published: (2024)
by: Cioflan, Cristian, et al.
Published: (2024)
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
by: Son, Donghyun, et al.
Published: (2025)
by: Son, Donghyun, et al.
Published: (2025)
Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping
by: Zheng, Xinzhe, et al.
Published: (2025)
by: Zheng, Xinzhe, et al.
Published: (2025)
GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units
by: Bouvier, Maxence, et al.
Published: (2025)
by: Bouvier, Maxence, et al.
Published: (2025)
Stella Nera: A Differentiable Maddness-Based Hardware Accelerator for Efficient Approximate Matrix Multiplication
by: Schönleber, Jannis, et al.
Published: (2023)
by: Schönleber, Jannis, et al.
Published: (2023)
Collage: Light-Weight Low-Precision Strategy for LLM Training
by: Yu, Tao, et al.
Published: (2024)
by: Yu, Tao, et al.
Published: (2024)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
by: Qiao, Ye, et al.
Published: (2026)
by: Qiao, Ye, et al.
Published: (2026)
GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration
by: Li, Yuhang, et al.
Published: (2025)
by: Li, Yuhang, et al.
Published: (2025)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
The Art of Beating the Odds with Predictor-Guided Random Design Space Exploration
by: Arnold, Felix, et al.
Published: (2025)
by: Arnold, Felix, et al.
Published: (2025)
Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method
by: Zhu, Qingcheng, et al.
Published: (2025)
by: Zhu, Qingcheng, et al.
Published: (2025)
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
Fast and Unified Path Gradient Estimators for Normalizing Flows
by: Vaitl, Lorenz, et al.
Published: (2024)
by: Vaitl, Lorenz, et al.
Published: (2024)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
by: Xu, Zifei, et al.
Published: (2024)
by: Xu, Zifei, et al.
Published: (2024)
AutoQRA: Joint Optimization of Mixed-Precision Quantization and Low-rank Adapters for Efficient LLM Fine-Tuning
by: Zhou, Changhai, et al.
Published: (2026)
by: Zhou, Changhai, et al.
Published: (2026)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
Neural Sinkhorn Gradient Flow
by: Zhu, Huminhao, et al.
Published: (2024)
by: Zhu, Huminhao, et al.
Published: (2024)
Sinkhorn-Drifting Generative Models
by: He, Ping, et al.
Published: (2026)
by: He, Ping, et al.
Published: (2026)
Federated Sinkhorn
by: Kulcsar, Jeremy, et al.
Published: (2025)
by: Kulcsar, Jeremy, et al.
Published: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
by: Mozaffari, Mohammad, et al.
Published: (2024)
by: Mozaffari, Mohammad, et al.
Published: (2024)
A singular Riemannian Geometry Approach to Deep Neural Networks III. Piecewise Differentiable Layers and Random Walks on $n$-dimensional Classes
by: Benfenati, Alessandro, et al.
Published: (2024)
by: Benfenati, Alessandro, et al.
Published: (2024)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
by: Shao, Yuantian, et al.
Published: (2025)
by: Shao, Yuantian, et al.
Published: (2025)
GSINA: Improving Subgraph Extraction for Graph Invariant Learning via Graph Sinkhorn Attention
by: Yan, Junchi, et al.
Published: (2024)
by: Yan, Junchi, et al.
Published: (2024)
Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart
by: Yu, Chengting, et al.
Published: (2024)
by: Yu, Chengting, et al.
Published: (2024)
Soft Quantization: Model Compression Via Weight Coupling
by: Bernstein, Daniel T., et al.
Published: (2026)
by: Bernstein, Daniel T., et al.
Published: (2026)
Confident Sinkhorn Allocation for Pseudo-Labeling
by: Nguyen, Vu, et al.
Published: (2022)
by: Nguyen, Vu, et al.
Published: (2022)
Coverage-Based Calibration for Post-Training Quantization via Weighted Set Cover over Outlier Channels
by: Shihab, Ibne Farabi, et al.
Published: (2026)
by: Shihab, Ibne Farabi, et al.
Published: (2026)
LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
by: Mirzaei, Amir Reza, et al.
Published: (2025)
by: Mirzaei, Amir Reza, et al.
Published: (2025)
Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks
by: Nishida, Keigo, et al.
Published: (2025)
by: Nishida, Keigo, et al.
Published: (2025)
Sinkhorn Distributionally Robust Optimization
by: Wang, Jie, et al.
Published: (2021)
by: Wang, Jie, et al.
Published: (2021)
Scalable Multivariate Fronthaul Quantization for Cell-Free Massive MIMO
by: Park, Sangwoo, et al.
Published: (2024)
by: Park, Sangwoo, et al.
Published: (2024)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
by: Zhou, Cyrus, et al.
Published: (2023)
by: Zhou, Cyrus, et al.
Published: (2023)
On Sinkhorn's Algorithm and Choice Modeling
by: Qu, Zhaonan, et al.
Published: (2023)
by: Qu, Zhaonan, et al.
Published: (2023)
Similar Items
-
TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025) -
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025) -
SSSD: Simply-Scalable Speculative Decoding
by: Marzollo, Michele, et al.
Published: (2024) -
Don't be so Stief! Learning KV Cache low-rank approximation over the Stiefel manifold
by: Benfenati, Luca, et al.
Published: (2026) -
Boosting keyword spotting through on-device learnable user speech characteristics
by: Cioflan, Cristian, et al.
Published: (2024)