Multi-Head Low-Rank Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Songtao, Peng, Hongwu, Zhang, Zhiwei, Chen, Zhengyu, Guo, Yue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
High-Layer Attention Pruning with Rescaling
von: Liu, Songtao, et al.
Veröffentlicht: (2025)
von: Liu, Songtao, et al.
Veröffentlicht: (2025)
Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
von: Diep, Nghiem T., et al.
Veröffentlicht: (2025)
von: Diep, Nghiem T., et al.
Veröffentlicht: (2025)
Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
von: Erden, Caner
Veröffentlicht: (2025)
von: Erden, Caner
Veröffentlicht: (2025)
Quantum Graph Attention Network: A Novel Quantum Multi-Head Attention Mechanism for Graph Learning
von: Ning, An, et al.
Veröffentlicht: (2025)
von: Ning, An, et al.
Veröffentlicht: (2025)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
Benign Overfitting in Single-Head Attention
von: Magen, Roey, et al.
Veröffentlicht: (2024)
von: Magen, Roey, et al.
Veröffentlicht: (2024)
MoH: Multi-Head Attention as Mixture-of-Head Attention
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Ravan: Multi-Head Low-Rank Adaptation for Federated Fine-Tuning
von: Raje, Arian, et al.
Veröffentlicht: (2025)
von: Raje, Arian, et al.
Veröffentlicht: (2025)
Adaptive Head Budgeting for Efficient Multi-Head Attention
von: Faye, Bilal, et al.
Veröffentlicht: (2026)
von: Faye, Bilal, et al.
Veröffentlicht: (2026)
Low-Rank Key Value Attention
von: O'Neill, James, et al.
Veröffentlicht: (2026)
von: O'Neill, James, et al.
Veröffentlicht: (2026)
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
Preference Optimization for Molecule Synthesis with Conditional Residual Energy-based Models
von: Liu, Songtao, et al.
Veröffentlicht: (2024)
von: Liu, Songtao, et al.
Veröffentlicht: (2024)
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
von: Cai, Tianle, et al.
Veröffentlicht: (2024)
von: Cai, Tianle, et al.
Veröffentlicht: (2024)
How Many Validation Labels Do You Need? Exploring the Design Space of Label-Efficient Model Ranking
von: Hu, Zhengyu, et al.
Veröffentlicht: (2023)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2023)
Simple yet Effective: Low-Rank Spatial Attention for Neural Operators
von: Yang, Zherui, et al.
Veröffentlicht: (2026)
von: Yang, Zherui, et al.
Veröffentlicht: (2026)
Memorization Capacity of Multi-Head Attention in Transformers
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2023)
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2023)
Multi-Head Residual-Gated DeepONet for Coherent Nonlinear Wave Dynamics
von: Fan, Zhiwei, et al.
Veröffentlicht: (2026)
von: Fan, Zhiwei, et al.
Veröffentlicht: (2026)
InRank: Incremental Low-Rank Learning
von: Zhao, Jiawei, et al.
Veröffentlicht: (2023)
von: Zhao, Jiawei, et al.
Veröffentlicht: (2023)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
von: Li, Cheng, et al.
Veröffentlicht: (2025)
von: Li, Cheng, et al.
Veröffentlicht: (2025)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
von: Gupta, Ahan, et al.
Veröffentlicht: (2023)
von: Gupta, Ahan, et al.
Veröffentlicht: (2023)
Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers
von: Chen, Anrui, et al.
Veröffentlicht: (2026)
von: Chen, Anrui, et al.
Veröffentlicht: (2026)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
Interleaved Head Attention
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2026)
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2026)
Boosting House Price Estimations with Multi-Head Gated Attention
von: Sellam, Zakaria Abdellah, et al.
Veröffentlicht: (2024)
von: Sellam, Zakaria Abdellah, et al.
Veröffentlicht: (2024)
Graph Contrastive Learning with Low-Rank Regularization and Low-Rank Attention for Noisy Node Classification
von: Wang, Yancheng, et al.
Veröffentlicht: (2024)
von: Wang, Yancheng, et al.
Veröffentlicht: (2024)
Provable Multi-Task Reinforcement Learning: A Representation Learning Framework with Low Rank Rewards
von: Guo, Yaoze, et al.
Veröffentlicht: (2026)
von: Guo, Yaoze, et al.
Veröffentlicht: (2026)
When and Why Grouping Attention Heads Accelerates Muon Optimization
von: Zhang, Hongtao, et al.
Veröffentlicht: (2026)
von: Zhang, Hongtao, et al.
Veröffentlicht: (2026)
GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
von: Liao, Xutao, et al.
Veröffentlicht: (2024)
von: Liao, Xutao, et al.
Veröffentlicht: (2024)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
Low Rank Multi-Dictionary Selection at Scale
von: Ma, Boya, et al.
Veröffentlicht: (2024)
von: Ma, Boya, et al.
Veröffentlicht: (2024)
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention
von: He, Jianliang, et al.
Veröffentlicht: (2025)
von: He, Jianliang, et al.
Veröffentlicht: (2025)
Beyond Parallelism: Synergistic Computational Graph Effects in Multi-Head Attention
von: Borde, Haitz Sáez de Ocáriz
Veröffentlicht: (2025)
von: Borde, Haitz Sáez de Ocáriz
Veröffentlicht: (2025)
MAGE: Multi-Head Attention Guided Embeddings for Low Resource Sentiment Classification
von: Vashisht, Varun, et al.
Veröffentlicht: (2025)
von: Vashisht, Varun, et al.
Veröffentlicht: (2025)
Geometric Analysis of Token Selection in Multi-Head Attention
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
von: Shahbazi, Ashkan, et al.
Veröffentlicht: (2025)
von: Shahbazi, Ashkan, et al.
Veröffentlicht: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
von: Xiao, Da, et al.
Veröffentlicht: (2024)
von: Xiao, Da, et al.
Veröffentlicht: (2024)
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
FedAH: Aggregated Head for Personalized Federated Learning
von: Zhou, Pengzhan, et al.
Veröffentlicht: (2024)
von: Zhou, Pengzhan, et al.
Veröffentlicht: (2024)
LoLA: Low-Rank Linear Attention With Sparse Caching
von: McDermott, Luke, et al.
Veröffentlicht: (2025)
von: McDermott, Luke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
High-Layer Attention Pruning with Rescaling
von: Liu, Songtao, et al.
Veröffentlicht: (2025) -
Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
von: Diep, Nghiem T., et al.
Veröffentlicht: (2025) -
Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
von: Erden, Caner
Veröffentlicht: (2025) -
Quantum Graph Attention Network: A Novel Quantum Multi-Head Attention Mechanism for Graph Learning
von: Ning, An, et al.
Veröffentlicht: (2025) -
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)