Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Erden, Caner |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multiscale Aggregated Hierarchical Attention (MAHA): A Game Theoretic and Optimization Driven Approach to Efficient Contextual Modeling in Large Language Models
von: Erden, Caner
Veröffentlicht: (2025)
von: Erden, Caner
Veröffentlicht: (2025)
Multi-Conditional Ranking with Large Language Models
von: Pezeshkpour, Pouya, et al.
Veröffentlicht: (2024)
von: Pezeshkpour, Pouya, et al.
Veröffentlicht: (2024)
Adaptive Feature-based Low-Rank Compression of Large Language Models via Bayesian Optimization
von: Ji, Yixin, et al.
Veröffentlicht: (2024)
von: Ji, Yixin, et al.
Veröffentlicht: (2024)
BoRA: Bayesian Hierarchical Low-Rank Adaption for Multi-Task Large Language Models
von: Eide, Simen, et al.
Veröffentlicht: (2024)
von: Eide, Simen, et al.
Veröffentlicht: (2024)
LCQ: Low-Rank Codebook based Quantization for Large Language Models
von: Cai, Wen-Pu, et al.
Veröffentlicht: (2024)
von: Cai, Wen-Pu, et al.
Veröffentlicht: (2024)
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)
LoLCATs: On Low-Rank Linearizing of Large Language Models
von: Zhang, Michael, et al.
Veröffentlicht: (2024)
von: Zhang, Michael, et al.
Veröffentlicht: (2024)
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing
von: Smith, James Seale, et al.
Veröffentlicht: (2025)
von: Smith, James Seale, et al.
Veröffentlicht: (2025)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
A Bayesian Interpretation of Adaptive Low-Rank Adaptation
von: Chen, Haolin, et al.
Veröffentlicht: (2024)
von: Chen, Haolin, et al.
Veröffentlicht: (2024)
LoLA: Low-Rank Linear Attention With Sparse Caching
von: McDermott, Luke, et al.
Veröffentlicht: (2025)
von: McDermott, Luke, et al.
Veröffentlicht: (2025)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
Found in the Middle: Permutation Self-Consistency Improves Listwise Ranking in Large Language Models
von: Tang, Raphael, et al.
Veröffentlicht: (2023)
von: Tang, Raphael, et al.
Veröffentlicht: (2023)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
von: Letzelter, Victor, et al.
Veröffentlicht: (2025)
von: Letzelter, Victor, et al.
Veröffentlicht: (2025)
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models
von: Shi, Haizhou, et al.
Veröffentlicht: (2024)
von: Shi, Haizhou, et al.
Veröffentlicht: (2024)
SoftLMs: Efficient Adaptive Low-Rank Approximation of Language Models using Soft-Thresholding Mechanism
von: Bhatnagar, Priyansh, et al.
Veröffentlicht: (2024)
von: Bhatnagar, Priyansh, et al.
Veröffentlicht: (2024)
Accelerating the Low-Rank Decomposed Models
von: Hajimolahoseini, Habib, et al.
Veröffentlicht: (2024)
von: Hajimolahoseini, Habib, et al.
Veröffentlicht: (2024)
SARA: Singular-Value Based Adaptive Low-Rank Adaption
von: Gu, Jihao, et al.
Veröffentlicht: (2024)
von: Gu, Jihao, et al.
Veröffentlicht: (2024)
Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
von: Roy, Debjyoti Saha, et al.
Veröffentlicht: (2024)
von: Roy, Debjyoti Saha, et al.
Veröffentlicht: (2024)
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
Demystifying Low-Rank Knowledge Distillation in Large Language Models: Convergence, Generalization, and Information-Theoretic Guarantees
von: Soarez, Alberlucia Rafael, et al.
Veröffentlicht: (2026)
von: Soarez, Alberlucia Rafael, et al.
Veröffentlicht: (2026)
Prompt-Dependent Ranking of Large Language Models with Uncertainty Quantification
von: Menendez, Angel Rodrigo Avelar, et al.
Veröffentlicht: (2026)
von: Menendez, Angel Rodrigo Avelar, et al.
Veröffentlicht: (2026)
Multi-Head Low-Rank Attention
von: Liu, Songtao, et al.
Veröffentlicht: (2026)
von: Liu, Songtao, et al.
Veröffentlicht: (2026)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
BA-LoRA: Bias-Alleviating Low-Rank Adaptation to Mitigate Catastrophic Inheritance in Large Language Models
von: Chang, Yupeng, et al.
Veröffentlicht: (2024)
von: Chang, Yupeng, et al.
Veröffentlicht: (2024)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
von: Gupta, Ahan, et al.
Veröffentlicht: (2023)
von: Gupta, Ahan, et al.
Veröffentlicht: (2023)
CoRA: Optimizing Low-Rank Adaptation with Common Subspace of Large Language Models
von: Xiao, Xiaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Xiaojun, et al.
Veröffentlicht: (2024)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
A3 : an Analytical Low-Rank Approximation Framework for Attention
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
LoRA+: Efficient Low Rank Adaptation of Large Models
von: Hayou, Soufiane, et al.
Veröffentlicht: (2024)
von: Hayou, Soufiane, et al.
Veröffentlicht: (2024)
AFLoRA: Adaptive Freezing of Low Rank Adaptation in Parameter Efficient Fine-Tuning of Large Models
von: Liu, Zeyu, et al.
Veröffentlicht: (2024)
von: Liu, Zeyu, et al.
Veröffentlicht: (2024)
GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
von: Liao, Xutao, et al.
Veröffentlicht: (2024)
von: Liao, Xutao, et al.
Veröffentlicht: (2024)
Sequences of Logits Reveal the Low Rank Structure of Language Models
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
Ranking Large Language Models without Ground Truth
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
Improving Transformers with Dynamically Composable Multi-Head Attention
von: Xiao, Da, et al.
Veröffentlicht: (2024)
von: Xiao, Da, et al.
Veröffentlicht: (2024)
Fast Forwarding Low-Rank Training
von: Rahamim, Adir, et al.
Veröffentlicht: (2024)
von: Rahamim, Adir, et al.
Veröffentlicht: (2024)
MTL-LoRA: Low-Rank Adaptation for Multi-Task Learning
von: Yang, Yaming, et al.
Veröffentlicht: (2024)
von: Yang, Yaming, et al.
Veröffentlicht: (2024)
Unsupervised Contrast-Consistent Ranking with Language Models
von: Stoehr, Niklas, et al.
Veröffentlicht: (2023)
von: Stoehr, Niklas, et al.
Veröffentlicht: (2023)
GoRA: Gradient-driven Adaptive Low Rank Adaptation
von: He, Haonan, et al.
Veröffentlicht: (2025)
von: He, Haonan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multiscale Aggregated Hierarchical Attention (MAHA): A Game Theoretic and Optimization Driven Approach to Efficient Contextual Modeling in Large Language Models
von: Erden, Caner
Veröffentlicht: (2025) -
Multi-Conditional Ranking with Large Language Models
von: Pezeshkpour, Pouya, et al.
Veröffentlicht: (2024) -
Adaptive Feature-based Low-Rank Compression of Large Language Models via Bayesian Optimization
von: Ji, Yixin, et al.
Veröffentlicht: (2024) -
BoRA: Bayesian Hierarchical Low-Rank Adaption for Multi-Task Large Language Models
von: Eide, Simen, et al.
Veröffentlicht: (2024) -
LCQ: Low-Rank Codebook based Quantization for Large Language Models
von: Cai, Wen-Pu, et al.
Veröffentlicht: (2024)