Different Prompts, Different Ranks: Prompt-aware Dynamic Rank Selection for SVD-based LLM Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Hengyi, Mi, Zhendong, Zhang, Grace Li, Huang, Shaoyi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-NAS: LLM-driven Hardware-Aware Neural Architecture Search
by: Zhu, Hengyi, et al.
Published: (2025)
by: Zhu, Hengyi, et al.
Published: (2025)
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
Layer-wise dynamic rank for compressing large language models
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Towards Fast LLM Fine-tuning through Zeroth-Order Optimization with Projected Gradient-Aligned Perturbations
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Effective MoE-based LLM Compression by Exploiting Heterogeneous Inter-Group Experts Routing Frequency and Information Density
by: Mi, Zhendong, et al.
Published: (2026)
by: Mi, Zhendong, et al.
Published: (2026)
Low-Rank Prehab: Preparing Neural Networks for SVD Compression
by: Qin, Haoran, et al.
Published: (2025)
by: Qin, Haoran, et al.
Published: (2025)
ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
by: Xv, Lin, et al.
Published: (2025)
by: Xv, Lin, et al.
Published: (2025)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
DipSVD: Dual-importance Protected SVD for Efficient LLM Compression
by: Ding, Xuan, et al.
Published: (2025)
by: Ding, Xuan, et al.
Published: (2025)
StealthRank: LLM Ranking Manipulation via Stealthy Prompt Optimization
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Unified Graph Prompt Learning via Low-Rank Graph Message Prompting
by: Wang, Beibei, et al.
Published: (2026)
by: Wang, Beibei, et al.
Published: (2026)
CoopetitiveV: Leveraging LLM-powered Coopetitive Multi-Agent Prompting for High-quality Verilog Generation
by: Mi, Zhendong, et al.
Published: (2024)
by: Mi, Zhendong, et al.
Published: (2024)
Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
by: Fang, Jiaxun, et al.
Published: (2025)
by: Fang, Jiaxun, et al.
Published: (2025)
Ranking-aware Reinforcement Learning for Ordinal Ranking
by: Hao, Aiming, et al.
Published: (2026)
by: Hao, Aiming, et al.
Published: (2026)
LASER: Low-Rank Activation SVD for Efficient Recursion
by: Çakar, Ege, et al.
Published: (2026)
by: Çakar, Ege, et al.
Published: (2026)
Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Low-Rank Compression of Language Models via Differentiable Rank Selection
by: Sundrani, Sidhant, et al.
Published: (2025)
by: Sundrani, Sidhant, et al.
Published: (2025)
Prompt Tuning Strikes Back: Customizing Foundation Models with Low-Rank Prompt Adaptation
by: Jain, Abhinav, et al.
Published: (2024)
by: Jain, Abhinav, et al.
Published: (2024)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
MoRA: On-the-fly Molecule-aware Low-Rank Adaptation Framework for LLM-based Multi-Modal Molecular Assistant
by: Yin, Tao, et al.
Published: (2025)
by: Yin, Tao, et al.
Published: (2025)
Operator SVD with Neural Networks via Nested Low-Rank Approximation
by: Ryu, J. Jon, et al.
Published: (2024)
by: Ryu, J. Jon, et al.
Published: (2024)
Prompt-Dependent Ranking of Large Language Models with Uncertainty Quantification
by: Menendez, Angel Rodrigo Avelar, et al.
Published: (2026)
by: Menendez, Angel Rodrigo Avelar, et al.
Published: (2026)
Hierarchical Sparse Plus Low Rank Compression of LLM
by: Kumar, Pawan, et al.
Published: (2025)
by: Kumar, Pawan, et al.
Published: (2025)
Efficient LLM Scheduling by Learning to Rank
by: Fu, Yichao, et al.
Published: (2024)
by: Fu, Yichao, et al.
Published: (2024)
Revisiting Weight Regularization for Low-Rank Continual Learning
by: Zheng, Yaoyue, et al.
Published: (2026)
by: Zheng, Yaoyue, et al.
Published: (2026)
AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression
by: Sinha, Atul Kumar, et al.
Published: (2026)
by: Sinha, Atul Kumar, et al.
Published: (2026)
Task-Driven Kernel Flows: Label Rank Compression and Laplacian Spectral Filtering
by: Li, Hongxi, et al.
Published: (2026)
by: Li, Hongxi, et al.
Published: (2026)
Beyond Uniform SVD:Dual-Level Optimization across Columns and Modules for LLM Compression
by: Xv, Lin, et al.
Published: (2025)
by: Xv, Lin, et al.
Published: (2025)
ProCut: LLM Prompt Compression via Attribution Estimation
by: Xu, Zhentao, et al.
Published: (2025)
by: Xu, Zhentao, et al.
Published: (2025)
LoRAP: Low-Rank Aggregation Prompting for Quantized Graph Neural Networks Training
by: Liu, Chenyu, et al.
Published: (2026)
by: Liu, Chenyu, et al.
Published: (2026)
Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
by: Solgi, Ryan, et al.
Published: (2025)
by: Solgi, Ryan, et al.
Published: (2025)
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting
by: Qin, Zhen, et al.
Published: (2023)
by: Qin, Zhen, et al.
Published: (2023)
AlphaRank: An Artificial Intelligence Approach for Ranking and Selection Problems
by: Zhou, Ruihan, et al.
Published: (2024)
by: Zhou, Ruihan, et al.
Published: (2024)
Additive Distributionally Robust Ranking and Selection
by: Li, Zaile, et al.
Published: (2025)
by: Li, Zaile, et al.
Published: (2025)
Accelerating Multi-Task Temporal Difference Learning under Low-Rank Representation
by: Bai, Yitao, et al.
Published: (2025)
by: Bai, Yitao, et al.
Published: (2025)
Prompt-SAW: Leveraging Relation-Aware Graphs for Textual Prompt Compression
by: Ali, Muhammad Asif, et al.
Published: (2024)
by: Ali, Muhammad Asif, et al.
Published: (2024)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
Similar Items
-
LLM-NAS: LLM-driven Hardware-Aware Neural Architecture Search
by: Zhu, Hengyi, et al.
Published: (2025) -
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026) -
Layer-wise dynamic rank for compressing large language models
by: Mi, Zhendong, et al.
Published: (2025) -
Towards Fast LLM Fine-tuning through Zeroth-Order Optimization with Projected Gradient-Aligned Perturbations
by: Mi, Zhendong, et al.
Published: (2025) -
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
by: Abbasi, Ali, et al.
Published: (2026)