Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhussip, Magauiya, Shopkhoev, Dmitriy, Ali, Ammar, Lefkimmiatis, Stamatios |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CoSpaDi: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
por: Makhov, Denis, et al.
Publicado: (2025)
por: Makhov, Denis, et al.
Publicado: (2025)
ROCKET: Rapid Optimization via Calibration-guided Knapsack Enhanced Truncation for Efficient Model Compression
por: Ali, Ammar, et al.
Publicado: (2026)
por: Ali, Ammar, et al.
Publicado: (2026)
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
por: Shopkhoev, Dmitriy, et al.
Publicado: (2025)
por: Shopkhoev, Dmitriy, et al.
Publicado: (2025)
Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
por: Mohammad, Baher, et al.
Publicado: (2025)
por: Mohammad, Baher, et al.
Publicado: (2025)
COMPOT: Calibration-Optimized Matrix Procrustes Orthogonalization for Transformers Compression
por: Makhov, Denis, et al.
Publicado: (2026)
por: Makhov, Denis, et al.
Publicado: (2026)
A Modular Conditional Diffusion Framework for Image Reconstruction
por: Zhussip, Magauiya, et al.
Publicado: (2024)
por: Zhussip, Magauiya, et al.
Publicado: (2024)
Beyond KV Caching: Shared Attention for Efficient LLMs
por: Liao, Bingli, et al.
Publicado: (2024)
por: Liao, Bingli, et al.
Publicado: (2024)
Advancing Arabic Reverse Dictionary Systems: A Transformer-Based Approach with Dataset Construction Guidelines
por: Sibaee, Serry, et al.
Publicado: (2025)
por: Sibaee, Serry, et al.
Publicado: (2025)
Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have)
por: Miyashita, Hisashi
Publicado: (2026)
por: Miyashita, Hisashi
Publicado: (2026)
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models
por: Liu, Xiaoze, et al.
Publicado: (2026)
por: Liu, Xiaoze, et al.
Publicado: (2026)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
por: Peng, Dan, et al.
Publicado: (2025)
por: Peng, Dan, et al.
Publicado: (2025)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
por: Yang, Yifei, et al.
Publicado: (2024)
por: Yang, Yifei, et al.
Publicado: (2024)
Your Transformer is Secretly Linear
por: Razzhigaev, Anton, et al.
Publicado: (2024)
por: Razzhigaev, Anton, et al.
Publicado: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
por: Deng, Yichuan, et al.
Publicado: (2024)
por: Deng, Yichuan, et al.
Publicado: (2024)
MoSLD: An Extremely Parameter-Efficient Mixture-of-Shared LoRAs for Multi-Task Learning
por: Zhao, Lulu, et al.
Publicado: (2024)
por: Zhao, Lulu, et al.
Publicado: (2024)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
por: Yang, Zhuonan, et al.
Publicado: (2026)
por: Yang, Zhuonan, et al.
Publicado: (2026)
Mixture of Attentions For Speculative Decoding
por: Zimmer, Matthieu, et al.
Publicado: (2024)
por: Zimmer, Matthieu, et al.
Publicado: (2024)
More Expressive Attention with Negative Weights
por: Lv, Ang, et al.
Publicado: (2024)
por: Lv, Ang, et al.
Publicado: (2024)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
por: Bhattacharjee, Arijit, et al.
Publicado: (2025)
por: Bhattacharjee, Arijit, et al.
Publicado: (2025)
Rethinking Parameter Sharing for LLM Fine-Tuning with Multiple LoRAs
por: Ban, Hao, et al.
Publicado: (2025)
por: Ban, Hao, et al.
Publicado: (2025)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
por: Chen, Yilong, et al.
Publicado: (2024)
por: Chen, Yilong, et al.
Publicado: (2024)
Features that Make a Difference: Leveraging Gradients for Improved Dictionary Learning
por: Olmo, Jeffrey, et al.
Publicado: (2024)
por: Olmo, Jeffrey, et al.
Publicado: (2024)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
por: Kyriakou, Athina, et al.
Publicado: (2026)
por: Kyriakou, Athina, et al.
Publicado: (2026)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
por: Lan, Michael, et al.
Publicado: (2023)
por: Lan, Michael, et al.
Publicado: (2023)
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
por: Karvonen, Adam, et al.
Publicado: (2024)
por: Karvonen, Adam, et al.
Publicado: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
por: Collins, Liam, et al.
Publicado: (2024)
por: Collins, Liam, et al.
Publicado: (2024)
How Private is Your Attention? Bridging Privacy with In-Context Learning
por: Bonnerjee, Soham, et al.
Publicado: (2025)
por: Bonnerjee, Soham, et al.
Publicado: (2025)
Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
por: Aden-Ali, Ishaq, et al.
Publicado: (2026)
por: Aden-Ali, Ishaq, et al.
Publicado: (2026)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
por: Benfeghoul, Martin, et al.
Publicado: (2025)
por: Benfeghoul, Martin, et al.
Publicado: (2025)
Selective Attention Improves Transformer
por: Leviathan, Yaniv, et al.
Publicado: (2024)
por: Leviathan, Yaniv, et al.
Publicado: (2024)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
por: Tian, Yuandong, et al.
Publicado: (2023)
por: Tian, Yuandong, et al.
Publicado: (2023)
Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
por: Liang, Yingyu, et al.
Publicado: (2024)
por: Liang, Yingyu, et al.
Publicado: (2024)
Hierarchical vs. Flat Iteration in Shared-Weight Transformers
por: Han, Sang-Il
Publicado: (2026)
por: Han, Sang-Il
Publicado: (2026)
Enhancing Hyperspace Analogue to Language (HAL) Representations via Attention-Based Pooling for Text Classification
por: Sakour, Ali, et al.
Publicado: (2026)
por: Sakour, Ali, et al.
Publicado: (2026)
Low-Cost Generation and Evaluation of Dictionary Example Sentences
por: Cai, Bill, et al.
Publicado: (2024)
por: Cai, Bill, et al.
Publicado: (2024)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
por: Leemann, Tobias, et al.
Publicado: (2024)
por: Leemann, Tobias, et al.
Publicado: (2024)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
por: Zhou, Sifan, et al.
Publicado: (2025)
por: Zhou, Sifan, et al.
Publicado: (2025)
KUET at StanceNakba Shared Task: StanceMoE: Mixture-of-Experts Architecture for Stance Detection
por: Shafi, Abdullah Al, et al.
Publicado: (2026)
por: Shafi, Abdullah Al, et al.
Publicado: (2026)
What Matters in Transformers? Not All Attention is Needed
por: He, Shwai, et al.
Publicado: (2024)
por: He, Shwai, et al.
Publicado: (2024)
Not All Synthetic Data Is Yours to Learn From
por: Alemohammad, Sina, et al.
Publicado: (2026)
por: Alemohammad, Sina, et al.
Publicado: (2026)
Ejemplares similares
-
CoSpaDi: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
por: Makhov, Denis, et al.
Publicado: (2025) -
ROCKET: Rapid Optimization via Calibration-guided Knapsack Enhanced Truncation for Efficient Model Compression
por: Ali, Ammar, et al.
Publicado: (2026) -
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
por: Shopkhoev, Dmitriy, et al.
Publicado: (2025) -
Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
por: Mohammad, Baher, et al.
Publicado: (2025) -
COMPOT: Calibration-Optimized Matrix Procrustes Orthogonalization for Transformers Compression
por: Makhov, Denis, et al.
Publicado: (2026)