Model Compression with Exact Budget Constraints via Riemannian Manifolds
Fuente:
arXiv
Saved in:
| Main Authors: | Helcig, Michael, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Statistically-Lossless Quantization of Large Language Models
by: Helcig, Michael, et al.
Published: (2026)
by: Helcig, Michael, et al.
Published: (2026)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
by: Dadgarnia, Alireza, et al.
Published: (2026)
by: Dadgarnia, Alireza, et al.
Published: (2026)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
by: Sieberling, Oliver, et al.
Published: (2024)
by: Sieberling, Oliver, et al.
Published: (2024)
FedCCL: Federated Clustered Continual Learning Framework for Privacy-focused Energy Forecasting
by: Helcig, Michael A., et al.
Published: (2025)
by: Helcig, Michael A., et al.
Published: (2025)
Extreme Compression of Large Language Models via Additive Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
by: Iofinova, Eugenia, et al.
Published: (2026)
by: Iofinova, Eugenia, et al.
Published: (2026)
Compression Scaling Laws:Unifying Sparsity and Quantization
by: Frantar, Elias, et al.
Published: (2025)
by: Frantar, Elias, et al.
Published: (2025)
Generalised Flow Maps for Few-Step Generative Modelling on Riemannian Manifolds
by: Davis, Oscar, et al.
Published: (2025)
by: Davis, Oscar, et al.
Published: (2025)
Apertus LLM Family Expansion via Distillation and Quantization
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
by: Schultheis, Erik, et al.
Published: (2025)
by: Schultheis, Erik, et al.
Published: (2025)
Riemannian Diffusion Models on General Manifolds via Physics-Informed Neural Networks
by: Ko, Gyeonghoon, et al.
Published: (2026)
by: Ko, Gyeonghoon, et al.
Published: (2026)
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
by: Kleinegger, Maximilian, et al.
Published: (2026)
by: Kleinegger, Maximilian, et al.
Published: (2026)
Powerset Convolutional Neural Networks
by: Wendler, Chris, et al.
Published: (2019)
by: Wendler, Chris, et al.
Published: (2019)
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
Error Feedback Can Accurately Compress Preconditioners
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
Conformal inference for regression on Riemannian Manifolds
by: Cholaquidis, Alejandro, et al.
Published: (2023)
by: Cholaquidis, Alejandro, et al.
Published: (2023)
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
Generative Modeling on Manifolds Through Mixture of Riemannian Diffusion Processes
by: Jo, Jaehyeong, et al.
Published: (2023)
by: Jo, Jaehyeong, et al.
Published: (2023)
Mirror Descent on Riemannian Manifolds
by: Jiang, Jiaxin, et al.
Published: (2026)
by: Jiang, Jiaxin, et al.
Published: (2026)
Simple Opinion Dynamics for No-Regret Learning
by: Lazarsfeld, John, et al.
Published: (2023)
by: Lazarsfeld, John, et al.
Published: (2023)
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
Dimensionality Reduction on Riemannian Manifolds in Data Analysis
by: Ichi, Alaa El, et al.
Published: (2026)
by: Ichi, Alaa El, et al.
Published: (2026)
Spiking Graph Neural Network on Riemannian Manifolds
by: Sun, Li, et al.
Published: (2024)
by: Sun, Li, et al.
Published: (2024)
Riemannian Optimization on Relaxed Indicator Matrix Manifold
by: Yuan, Jinghui, et al.
Published: (2025)
by: Yuan, Jinghui, et al.
Published: (2025)
Efficient Data Selection at Scale via Influence Distillation
by: Nikdan, Mahdi, et al.
Published: (2025)
by: Nikdan, Mahdi, et al.
Published: (2025)
Sharpness-Aware Teleportation on Riemannian Manifolds
by: Truong, Tuan, et al.
Published: (2023)
by: Truong, Tuan, et al.
Published: (2023)
Efficient Sampling on Riemannian Manifolds via Langevin MCMC
by: Cheng, Xiang, et al.
Published: (2024)
by: Cheng, Xiang, et al.
Published: (2024)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Communication-Efficient Federated Learning With Data and Client Heterogeneity
by: Zakerinia, Hossein, et al.
Published: (2022)
by: Zakerinia, Hossein, et al.
Published: (2022)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
by: Nikdan, Mahdi, et al.
Published: (2024)
by: Nikdan, Mahdi, et al.
Published: (2024)
GeoDynamics: A Geometric State-Space Neural Network for Understanding Brain Dynamics on Riemannian Manifolds
by: Dan, Tingting, et al.
Published: (2026)
by: Dan, Tingting, et al.
Published: (2026)
Dual Riemannian Newton Method on Statistical Manifolds
by: Zhou, Derun, et al.
Published: (2025)
by: Zhou, Derun, et al.
Published: (2025)
Riemannian MeanFlow for One-Step Generation on Manifolds
by: Zhong, Zichen, et al.
Published: (2026)
by: Zhong, Zichen, et al.
Published: (2026)
ManifoldFormer: Geometric Deep Learning for Neural Dynamics on Riemannian Manifolds
by: Fu, Yihang, et al.
Published: (2025)
by: Fu, Yihang, et al.
Published: (2025)
MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
by: Frantar, Elias, et al.
Published: (2024)
by: Frantar, Elias, et al.
Published: (2024)
Riemannian Optimization for LoRA on the Stiefel Manifold
by: Park, Juneyoung, et al.
Published: (2025)
by: Park, Juneyoung, et al.
Published: (2025)
A Framework for Bilevel Optimization on Riemannian Manifolds
by: Han, Andi, et al.
Published: (2024)
by: Han, Andi, et al.
Published: (2024)
Towards Robust Scaling Laws for Optimizers
by: Volkova, Alexandra, et al.
Published: (2026)
by: Volkova, Alexandra, et al.
Published: (2026)
Similar Items
-
Statistically-Lossless Quantization of Large Language Models
by: Helcig, Michael, et al.
Published: (2026) -
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
by: Dadgarnia, Alireza, et al.
Published: (2026) -
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
by: Sieberling, Oliver, et al.
Published: (2024) -
FedCCL: Federated Clustered Continual Learning Framework for Privacy-focused Energy Forecasting
by: Helcig, Michael A., et al.
Published: (2025) -
Extreme Compression of Large Language Models via Additive Quantization
by: Egiazarian, Vage, et al.
Published: (2024)