ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Xv, Lin, Gao, Jingsheng, Gao, Xian, Liu, Ting, Fu, Yuzhuo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Uniform SVD:Dual-Level Optimization across Columns and Modules for LLM Compression
by: Xv, Lin, et al.
Published: (2025)
by: Xv, Lin, et al.
Published: (2025)
AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression
by: Sinha, Atul Kumar, et al.
Published: (2026)
by: Sinha, Atul Kumar, et al.
Published: (2026)
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
DipSVD: Dual-importance Protected SVD for Efficient LLM Compression
by: Ding, Xuan, et al.
Published: (2025)
by: Ding, Xuan, et al.
Published: (2025)
Low-Rank Prehab: Preparing Neural Networks for SVD Compression
by: Qin, Haoran, et al.
Published: (2025)
by: Qin, Haoran, et al.
Published: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
LASER: Low-Rank Activation SVD for Efficient Recursion
by: Çakar, Ege, et al.
Published: (2026)
by: Çakar, Ege, et al.
Published: (2026)
CALR: Corrective Adaptive Low-Rank Decomposition for Efficient Large Language Model Layer Compression
by: Kautsar, Muchammad Daniyal, et al.
Published: (2025)
by: Kautsar, Muchammad Daniyal, et al.
Published: (2025)
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
Different Prompts, Different Ranks: Prompt-aware Dynamic Rank Selection for SVD-based LLM Compression
by: Zhu, Hengyi, et al.
Published: (2026)
by: Zhu, Hengyi, et al.
Published: (2026)
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models
by: Chekalina, Viktoriia, et al.
Published: (2025)
by: Chekalina, Viktoriia, et al.
Published: (2025)
MGAA: Multi-Granular Adaptive Allocation fof Low-Rank Compression of LLMs
by: Li, Guangyan, et al.
Published: (2025)
by: Li, Guangyan, et al.
Published: (2025)
Adaptive Feature-based Low-Rank Compression of Large Language Models via Bayesian Optimization
by: Ji, Yixin, et al.
Published: (2024)
by: Ji, Yixin, et al.
Published: (2024)
Adaptive Rank Allocation for Federated Parameter-Efficient Fine-Tuning of Language Models
by: Wu, Fei, et al.
Published: (2025)
by: Wu, Fei, et al.
Published: (2025)
ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews
by: Gao, Xian, et al.
Published: (2025)
by: Gao, Xian, et al.
Published: (2025)
MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
by: Gao, Xian, et al.
Published: (2025)
by: Gao, Xian, et al.
Published: (2025)
Large Language Model Compression with Global Rank and Sparsity Optimization
by: Zhou, Changhai, et al.
Published: (2025)
by: Zhou, Changhai, et al.
Published: (2025)
BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
by: Xu, Peng, et al.
Published: (2024)
by: Xu, Peng, et al.
Published: (2024)
Adaptive Rank Allocation: Speeding Up Modern Transformers with RaNA Adapters
by: Garcia, Roberto, et al.
Published: (2025)
by: Garcia, Roberto, et al.
Published: (2025)
EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
by: Yu, Zhongzhi, et al.
Published: (2024)
by: Yu, Zhongzhi, et al.
Published: (2024)
KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
by: Lesens, Damien, et al.
Published: (2025)
by: Lesens, Damien, et al.
Published: (2025)
Capability-Guided Compression: Toward Interpretability-Aware Budget Allocation for Large Language Models
by: Gupta, Rishaank
Published: (2026)
by: Gupta, Rishaank
Published: (2026)
EIAD: Explainable Industrial Anomaly Detection Via Multi-Modal Large Language Models
by: Zhang, Zongyun, et al.
Published: (2025)
by: Zhang, Zongyun, et al.
Published: (2025)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
by: Zhang, Chao, et al.
Published: (2025)
by: Zhang, Chao, et al.
Published: (2025)
Lillama: Large Language Models Compression via Low-Rank Feature Distillation
by: Sy, Yaya, et al.
Published: (2024)
by: Sy, Yaya, et al.
Published: (2024)
LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models
by: Wang, Haiyu, et al.
Published: (2026)
by: Wang, Haiyu, et al.
Published: (2026)
Dynamic Low-Rank Sparse Adaptation for Large Language Models
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
Low-Rank Compression of Language Models via Differentiable Rank Selection
by: Sundrani, Sidhant, et al.
Published: (2025)
by: Sundrani, Sidhant, et al.
Published: (2025)
Model Compression and Efficient Inference for Large Language Models: A Survey
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
Goal-Guided Efficient Exploration via Large Language Model in Reinforcement Learning
by: Qi, Yajie, et al.
Published: (2025)
by: Qi, Yajie, et al.
Published: (2025)
Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Differentially Private Low-Rank Adaptation of Large Language Model Using Federated Learning
by: Liu, Xiao-Yang, et al.
Published: (2023)
by: Liu, Xiao-Yang, et al.
Published: (2023)
CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor Optimization
by: Yang, Zi, et al.
Published: (2024)
by: Yang, Zi, et al.
Published: (2024)
Operator SVD with Neural Networks via Nested Low-Rank Approximation
by: Ryu, J. Jon, et al.
Published: (2024)
by: Ryu, J. Jon, et al.
Published: (2024)
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing
by: Smith, James Seale, et al.
Published: (2025)
by: Smith, James Seale, et al.
Published: (2025)
Efficient Annotator Reliability Assessment with EffiARA
by: Cook, Owen, et al.
Published: (2025)
by: Cook, Owen, et al.
Published: (2025)
ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer
by: Deng, Chunyuan, et al.
Published: (2026)
by: Deng, Chunyuan, et al.
Published: (2026)
SVD-NO: Learning PDE Solution Operators with SVD Integral Kernels
by: Koren, Noam, et al.
Published: (2025)
by: Koren, Noam, et al.
Published: (2025)
Similar Items
-
Beyond Uniform SVD:Dual-Level Optimization across Columns and Modules for LLM Compression
by: Xv, Lin, et al.
Published: (2025) -
AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression
by: Sinha, Atul Kumar, et al.
Published: (2026) -
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026) -
DipSVD: Dual-importance Protected SVD for Efficient LLM Compression
by: Ding, Xuan, et al.
Published: (2025) -
Low-Rank Prehab: Preparing Neural Networks for SVD Compression
by: Qin, Haoran, et al.
Published: (2025)