Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Abbasi, Ali, Thrash, Chayne, Qin, Haoran, Sharma, Shansita, Seifi, Sepehr, Kolouri, Soheil |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Low-Rank Prehab: Preparing Neural Networks for SVD Compression
by: Qin, Haoran, et al.
Published: (2025)
by: Qin, Haoran, et al.
Published: (2025)
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
by: Thrash, Chayne, et al.
Published: (2026)
by: Thrash, Chayne, et al.
Published: (2026)
MCNC: Manifold-Constrained Reparameterization for Neural Compression
by: Thrash, Chayne, et al.
Published: (2024)
by: Thrash, Chayne, et al.
Published: (2024)
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
by: Shahbazi, Ashkan, et al.
Published: (2025)
by: Shahbazi, Ashkan, et al.
Published: (2025)
EMPEROR: Efficient Moment-Preserving Representation of Distributions
by: Liu, Xinran, et al.
Published: (2025)
by: Liu, Xinran, et al.
Published: (2025)
OT-MeanFlow3D: Bridging Optimal Transport and Meanflow for Efficient 3D Point Cloud Generation
by: Akbari, Elaheh, et al.
Published: (2025)
by: Akbari, Elaheh, et al.
Published: (2025)
One Category One Prompt: Dataset Distillation using Diffusion Models
by: Abbasi, Ali, et al.
Published: (2024)
by: Abbasi, Ali, et al.
Published: (2024)
LASER: Low-Rank Activation SVD for Efficient Recursion
by: Çakar, Ege, et al.
Published: (2026)
by: Çakar, Ege, et al.
Published: (2026)
Different Prompts, Different Ranks: Prompt-aware Dynamic Rank Selection for SVD-based LLM Compression
by: Zhu, Hengyi, et al.
Published: (2026)
by: Zhu, Hengyi, et al.
Published: (2026)
Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
Physics informed cell representations for variational formulation of multiscale problems
by: Gao, Yuxiang, et al.
Published: (2024)
by: Gao, Yuxiang, et al.
Published: (2024)
DipSVD: Dual-importance Protected SVD for Efficient LLM Compression
by: Ding, Xuan, et al.
Published: (2025)
by: Ding, Xuan, et al.
Published: (2025)
Efficient Federated Low Rank Matrix Completion
by: Abbasi, Ahmed Ali, et al.
Published: (2024)
by: Abbasi, Ahmed Ali, et al.
Published: (2024)
ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
by: Xv, Lin, et al.
Published: (2025)
by: Xv, Lin, et al.
Published: (2025)
Vector-Quantized Soft Label Compression for Dataset Distillation
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression
by: Sinha, Atul Kumar, et al.
Published: (2026)
by: Sinha, Atul Kumar, et al.
Published: (2026)
LUNA: Linear Universal Neural Attention with Generalization Guarantees
by: Shahbazi, Ashkan, et al.
Published: (2025)
by: Shahbazi, Ashkan, et al.
Published: (2025)
Operator SVD with Neural Networks via Nested Low-Rank Approximation
by: Ryu, J. Jon, et al.
Published: (2024)
by: Ryu, J. Jon, et al.
Published: (2024)
Sinkhorn-Drifting Generative Models
by: He, Ping, et al.
Published: (2026)
by: He, Ping, et al.
Published: (2026)
Min Generalized Sliced Gromov Wasserstein: A Scalable Path to Gromov Wasserstein
by: Shahbazi, Ashkan, et al.
Published: (2026)
by: Shahbazi, Ashkan, et al.
Published: (2026)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
Hierarchical Sparse Plus Low Rank Compression of LLM
by: Kumar, Pawan, et al.
Published: (2025)
by: Kumar, Pawan, et al.
Published: (2025)
Beyond Uniform SVD:Dual-Level Optimization across Columns and Modules for LLM Compression
by: Xv, Lin, et al.
Published: (2025)
by: Xv, Lin, et al.
Published: (2025)
SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models
by: Hong, Chengjie, et al.
Published: (2026)
by: Hong, Chengjie, et al.
Published: (2026)
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Fused Partial Gromov-Wasserstein for Structured Objects
by: Bai, Yikun, et al.
Published: (2025)
by: Bai, Yikun, et al.
Published: (2025)
ESPFormer: Doubly-Stochastic Attention with Expected Sliced Transport Plans
by: Shahbazi, Ashkan, et al.
Published: (2025)
by: Shahbazi, Ashkan, et al.
Published: (2025)
Linear Partial Gromov-Wasserstein Embedding
by: Bai, Yikun, et al.
Published: (2024)
by: Bai, Yikun, et al.
Published: (2024)
Constrained Sliced Wasserstein Embedding
by: NaderiAlizadeh, Navid, et al.
Published: (2025)
by: NaderiAlizadeh, Navid, et al.
Published: (2025)
Neural-Augmented Kelvinlet for Real-Time Soft Tissue Deformation Modeling
by: Shahbazi, Ashkan, et al.
Published: (2025)
by: Shahbazi, Ashkan, et al.
Published: (2025)
Equivariant vs. Invariant Layers: A Comparison of Backbone and Pooling for Point Cloud Classification
by: Kothapalli, Abihith, et al.
Published: (2023)
by: Kothapalli, Abihith, et al.
Published: (2023)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
Adaptive Pruning with Module Robustness Sensitivity: Balancing Compression and Robustness
by: Bai, Lincen, et al.
Published: (2024)
by: Bai, Lincen, et al.
Published: (2024)
Partial Gromov-Wasserstein Metric
by: Bai, Yikun, et al.
Published: (2024)
by: Bai, Yikun, et al.
Published: (2024)
KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
by: Lesens, Damien, et al.
Published: (2025)
by: Lesens, Damien, et al.
Published: (2025)
Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
by: Solgi, Ryan, et al.
Published: (2025)
by: Solgi, Ryan, et al.
Published: (2025)
DeltaLLM: Compress LLMs with Low-Rank Deltas between Shared Weights
by: Mikaelyan, Liana, et al.
Published: (2025)
by: Mikaelyan, Liana, et al.
Published: (2025)
Statistical Context Detection for Deep Lifelong Reinforcement Learning
by: Dick, Jeffery, et al.
Published: (2024)
by: Dick, Jeffery, et al.
Published: (2024)
SVD Contextual Sparsity Predictors for Fast LLM Inference
by: Serbin, Georgii, et al.
Published: (2026)
by: Serbin, Georgii, et al.
Published: (2026)
Similar Items
-
Low-Rank Prehab: Preparing Neural Networks for SVD Compression
by: Qin, Haoran, et al.
Published: (2025) -
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026) -
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
by: Thrash, Chayne, et al.
Published: (2026) -
MCNC: Manifold-Constrained Reparameterization for Neural Compression
by: Thrash, Chayne, et al.
Published: (2024) -
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
by: Shahbazi, Ashkan, et al.
Published: (2025)