LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Haiyu, Wang, Yutong, Li, Leshu, Ren, Yihui, Zhang, Sai Qian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models
von: Wang, Haiyu, et al.
Veröffentlicht: (2026)
von: Wang, Haiyu, et al.
Veröffentlicht: (2026)
QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
LASER: Low-Rank Activation SVD for Efficient Recursion
von: Çakar, Ege, et al.
Veröffentlicht: (2026)
von: Çakar, Ege, et al.
Veröffentlicht: (2026)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
Compressing Large Language Models using Low Rank and Low Precision Decomposition
von: Saha, Rajarshi, et al.
Veröffentlicht: (2024)
von: Saha, Rajarshi, et al.
Veröffentlicht: (2024)
EDoRA: Efficient Weight-Decomposed Low-Rank Adaptation via Singular Value Decomposition
von: Nasiri, Hamid, et al.
Veröffentlicht: (2025)
von: Nasiri, Hamid, et al.
Veröffentlicht: (2025)
Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models
von: Zhang, Longteng, et al.
Veröffentlicht: (2026)
von: Zhang, Longteng, et al.
Veröffentlicht: (2026)
PLAN: Proactive Low-Rank Allocation for Continual Learning
von: Wang, Xiequn, et al.
Veröffentlicht: (2025)
von: Wang, Xiequn, et al.
Veröffentlicht: (2025)
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
LASER: Attention with Exponential Transformation
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2024)
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2024)
CALR: Corrective Adaptive Low-Rank Decomposition for Efficient Large Language Model Layer Compression
von: Kautsar, Muchammad Daniyal, et al.
Veröffentlicht: (2025)
von: Kautsar, Muchammad Daniyal, et al.
Veröffentlicht: (2025)
SARA: Singular-Value Based Adaptive Low-Rank Adaption
von: Gu, Jihao, et al.
Veröffentlicht: (2024)
von: Gu, Jihao, et al.
Veröffentlicht: (2024)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
Importance-Guided Basis Selection for Low-Rank Decomposition of Large Language Models
von: Asante, Daniel Agyei, et al.
Veröffentlicht: (2026)
von: Asante, Daniel Agyei, et al.
Veröffentlicht: (2026)
ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
von: Xv, Lin, et al.
Veröffentlicht: (2025)
von: Xv, Lin, et al.
Veröffentlicht: (2025)
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
Maestro: Uncovering Low-Rank Structures via Trainable Decomposition
von: Horvath, Samuel, et al.
Veröffentlicht: (2023)
von: Horvath, Samuel, et al.
Veröffentlicht: (2023)
FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment
von: Zaccone, Riccardo, et al.
Veröffentlicht: (2026)
von: Zaccone, Riccardo, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning with Generative Trajectory Policies
von: Feng, Xinsong, et al.
Veröffentlicht: (2025)
von: Feng, Xinsong, et al.
Veröffentlicht: (2025)
On Catastrophic Forgetting in Low-Rank Decomposition-Based Parameter-Efficient Fine-Tuning
von: Ahmad, Muhammad, et al.
Veröffentlicht: (2026)
von: Ahmad, Muhammad, et al.
Veröffentlicht: (2026)
The Panaceas for Improving Low-Rank Decomposition in Communication-Efficient Federated Learning
von: Li, Shiwei, et al.
Veröffentlicht: (2025)
von: Li, Shiwei, et al.
Veröffentlicht: (2025)
AFLoRA: Adaptive Federated Fine-Tuning of Large Language Models with Resource-Aware Low-Rank Adaption
von: Zhou, Yajie, et al.
Veröffentlicht: (2025)
von: Zhou, Yajie, et al.
Veröffentlicht: (2025)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
von: Lee, Seoungsub, et al.
Veröffentlicht: (2026)
von: Lee, Seoungsub, et al.
Veröffentlicht: (2026)
TalkLoRA: Communication-Aware Mixture of Low-Rank Adaptation for Large Language Models
von: Mu, Lin, et al.
Veröffentlicht: (2026)
von: Mu, Lin, et al.
Veröffentlicht: (2026)
MAP: Revisiting Weight Decomposition for Low-Rank Adaptation
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
CTR-LoRA: Curvature-Aware and Trust-Region Guided Low-Rank Adaptation for Large Language Models
von: Wang, Zhuxuanzi, et al.
Veröffentlicht: (2025)
von: Wang, Zhuxuanzi, et al.
Veröffentlicht: (2025)
SEMU: Singular Value Decomposition for Efficient Machine Unlearning
von: Sendera, Marcin, et al.
Veröffentlicht: (2025)
von: Sendera, Marcin, et al.
Veröffentlicht: (2025)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models
von: Ghiasvand, Sajjad, et al.
Veröffentlicht: (2025)
von: Ghiasvand, Sajjad, et al.
Veröffentlicht: (2025)
C-LoRA: Contextual Low-Rank Adaptation for Uncertainty Estimation in Large Language Models
von: Rahmati, Amir Hossein, et al.
Veröffentlicht: (2025)
von: Rahmati, Amir Hossein, et al.
Veröffentlicht: (2025)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
Adaptive Rank Allocation for Federated Parameter-Efficient Fine-Tuning of Language Models
von: Wu, Fei, et al.
Veröffentlicht: (2025)
von: Wu, Fei, et al.
Veröffentlicht: (2025)
Singular Value Decomposition on Kronecker Adaptation for Large Language Model
von: Chong, Yee Hin, et al.
Veröffentlicht: (2025)
von: Chong, Yee Hin, et al.
Veröffentlicht: (2025)
Communication-Efficient Personalized Federal Graph Learning via Low-Rank Decomposition
von: Liu, Ruyue, et al.
Veröffentlicht: (2024)
von: Liu, Ruyue, et al.
Veröffentlicht: (2024)
Compressed BC-LISTA via Low-Rank Convolutional Decomposition
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
FedSLoP: Memory-Efficient Federated Learning with Low-Rank Gradient Projection
von: He, Yutong, et al.
Veröffentlicht: (2026)
von: He, Yutong, et al.
Veröffentlicht: (2026)
Context Features Are Cheap: Rank-Aware Decomposition for Efficient Feature Interaction in Recommender Systems
von: Tkach, Yevgeny
Veröffentlicht: (2026)
von: Tkach, Yevgeny
Veröffentlicht: (2026)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
Optimal Brain Decomposition for Accurate LLM Low-Rank Approximation
von: Li, Yuhang, et al.
Veröffentlicht: (2026)
von: Li, Yuhang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models
von: Wang, Haiyu, et al.
Veröffentlicht: (2026) -
QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
von: Wang, Yutong, et al.
Veröffentlicht: (2025) -
LASER: Low-Rank Activation SVD for Efficient Recursion
von: Çakar, Ege, et al.
Veröffentlicht: (2026) -
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
von: Yu, Fengze, et al.
Veröffentlicht: (2025) -
Compressing Large Language Models using Low Rank and Low Precision Decomposition
von: Saha, Rajarshi, et al.
Veröffentlicht: (2024)