Low-Rank Prehab: Preparing Neural Networks for SVD Compression
Fuente:
arXiv
Guardado en:
| Autores principales: | Qin, Haoran, Sharma, Shansita, Abbasi, Ali, Thrash, Chayne, Kolouri, Soheil |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
por: Abbasi, Ali, et al.
Publicado: (2026)
por: Abbasi, Ali, et al.
Publicado: (2026)
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
por: Abbasi, Ali, et al.
Publicado: (2026)
por: Abbasi, Ali, et al.
Publicado: (2026)
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
por: Thrash, Chayne, et al.
Publicado: (2026)
por: Thrash, Chayne, et al.
Publicado: (2026)
MCNC: Manifold-Constrained Reparameterization for Neural Compression
por: Thrash, Chayne, et al.
Publicado: (2024)
por: Thrash, Chayne, et al.
Publicado: (2024)
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
EMPEROR: Efficient Moment-Preserving Representation of Distributions
por: Liu, Xinran, et al.
Publicado: (2025)
por: Liu, Xinran, et al.
Publicado: (2025)
OT-MeanFlow3D: Bridging Optimal Transport and Meanflow for Efficient 3D Point Cloud Generation
por: Akbari, Elaheh, et al.
Publicado: (2025)
por: Akbari, Elaheh, et al.
Publicado: (2025)
One Category One Prompt: Dataset Distillation using Diffusion Models
por: Abbasi, Ali, et al.
Publicado: (2024)
por: Abbasi, Ali, et al.
Publicado: (2024)
LASER: Low-Rank Activation SVD for Efficient Recursion
por: Çakar, Ege, et al.
Publicado: (2026)
por: Çakar, Ege, et al.
Publicado: (2026)
Operator SVD with Neural Networks via Nested Low-Rank Approximation
por: Ryu, J. Jon, et al.
Publicado: (2024)
por: Ryu, J. Jon, et al.
Publicado: (2024)
Physics informed cell representations for variational formulation of multiscale problems
por: Gao, Yuxiang, et al.
Publicado: (2024)
por: Gao, Yuxiang, et al.
Publicado: (2024)
LUNA: Linear Universal Neural Attention with Generalization Guarantees
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
Efficient Federated Low Rank Matrix Completion
por: Abbasi, Ahmed Ali, et al.
Publicado: (2024)
por: Abbasi, Ahmed Ali, et al.
Publicado: (2024)
Low-Rank Matrix Approximation for Neural Network Compression
por: Cherukuri, Kalyan, et al.
Publicado: (2025)
por: Cherukuri, Kalyan, et al.
Publicado: (2025)
ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
por: Xv, Lin, et al.
Publicado: (2025)
por: Xv, Lin, et al.
Publicado: (2025)
Different Prompts, Different Ranks: Prompt-aware Dynamic Rank Selection for SVD-based LLM Compression
por: Zhu, Hengyi, et al.
Publicado: (2026)
por: Zhu, Hengyi, et al.
Publicado: (2026)
Vector-Quantized Soft Label Compression for Dataset Distillation
por: Abbasi, Ali, et al.
Publicado: (2026)
por: Abbasi, Ali, et al.
Publicado: (2026)
Theoretical Guarantees for Low-Rank Compression of Deep Neural Networks
por: Zhang, Shihao, et al.
Publicado: (2025)
por: Zhang, Shihao, et al.
Publicado: (2025)
Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
por: Wang, Qinsi, et al.
Publicado: (2025)
por: Wang, Qinsi, et al.
Publicado: (2025)
AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression
por: Sinha, Atul Kumar, et al.
Publicado: (2026)
por: Sinha, Atul Kumar, et al.
Publicado: (2026)
Sinkhorn-Drifting Generative Models
por: He, Ping, et al.
Publicado: (2026)
por: He, Ping, et al.
Publicado: (2026)
DipSVD: Dual-importance Protected SVD for Efficient LLM Compression
por: Ding, Xuan, et al.
Publicado: (2025)
por: Ding, Xuan, et al.
Publicado: (2025)
Min Generalized Sliced Gromov Wasserstein: A Scalable Path to Gromov Wasserstein
por: Shahbazi, Ashkan, et al.
Publicado: (2026)
por: Shahbazi, Ashkan, et al.
Publicado: (2026)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
por: Shao, Zishan, et al.
Publicado: (2025)
por: Shao, Zishan, et al.
Publicado: (2025)
Dynamical Low-Rank Compression of Neural Networks with Robustness under Adversarial Attacks
por: Schotthöfer, Steffen, et al.
Publicado: (2025)
por: Schotthöfer, Steffen, et al.
Publicado: (2025)
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
por: Patel, Niket, et al.
Publicado: (2024)
por: Patel, Niket, et al.
Publicado: (2024)
Fused Partial Gromov-Wasserstein for Structured Objects
por: Bai, Yikun, et al.
Publicado: (2025)
por: Bai, Yikun, et al.
Publicado: (2025)
ESPFormer: Doubly-Stochastic Attention with Expected Sliced Transport Plans
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
Linear Partial Gromov-Wasserstein Embedding
por: Bai, Yikun, et al.
Publicado: (2024)
por: Bai, Yikun, et al.
Publicado: (2024)
Constrained Sliced Wasserstein Embedding
por: NaderiAlizadeh, Navid, et al.
Publicado: (2025)
por: NaderiAlizadeh, Navid, et al.
Publicado: (2025)
Equivariant vs. Invariant Layers: A Comparison of Backbone and Pooling for Point Cloud Classification
por: Kothapalli, Abihith, et al.
Publicado: (2023)
por: Kothapalli, Abihith, et al.
Publicado: (2023)
HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
por: Xiao, Jinqi, et al.
Publicado: (2023)
por: Xiao, Jinqi, et al.
Publicado: (2023)
On Generalization Bounds for Neural Networks with Low Rank Layers
por: Pinto, Andrea, et al.
Publicado: (2024)
por: Pinto, Andrea, et al.
Publicado: (2024)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
por: Wu, Wenhao, et al.
Publicado: (2026)
por: Wu, Wenhao, et al.
Publicado: (2026)
Partial Gromov-Wasserstein Metric
por: Bai, Yikun, et al.
Publicado: (2024)
por: Bai, Yikun, et al.
Publicado: (2024)
KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
por: Lesens, Damien, et al.
Publicado: (2025)
por: Lesens, Damien, et al.
Publicado: (2025)
Unified Framework for Pre-trained Neural Network Compression via Decomposition and Optimized Rank Selection
por: Aghababaei-Harandi, Ali, et al.
Publicado: (2024)
por: Aghababaei-Harandi, Ali, et al.
Publicado: (2024)
Statistical Context Detection for Deep Lifelong Reinforcement Learning
por: Dick, Jeffery, et al.
Publicado: (2024)
por: Dick, Jeffery, et al.
Publicado: (2024)
Low-Rank Adapters Meet Neural Architecture Search for LLM Compression
por: Muñoz, J. Pablo, et al.
Publicado: (2025)
por: Muñoz, J. Pablo, et al.
Publicado: (2025)
Low-Rank Tensor Decompositions for the Theory of Neural Networks
por: Borsoi, Ricardo, et al.
Publicado: (2025)
por: Borsoi, Ricardo, et al.
Publicado: (2025)
Ejemplares similares
-
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
por: Abbasi, Ali, et al.
Publicado: (2026) -
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
por: Abbasi, Ali, et al.
Publicado: (2026) -
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
por: Thrash, Chayne, et al.
Publicado: (2026) -
MCNC: Manifold-Constrained Reparameterization for Neural Compression
por: Thrash, Chayne, et al.
Publicado: (2024) -
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
por: Shahbazi, Ashkan, et al.
Publicado: (2025)