Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Karbevski, Marko, Mijoski, Antonij |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Can an MLP Absorb Its Own Skip Connection?
par: Mijoski, Antonij, et autres
Publié: (2026)
par: Mijoski, Antonij, et autres
Publié: (2026)
Beyond Linearity in Attention Projections: The Case for Nonlinear Queries
par: Karbevski, Marko
Publié: (2026)
par: Karbevski, Marko
Publié: (2026)
Attention is All You Need Until You Need Retention
par: Yaslioglu, M. Murat
Publié: (2025)
par: Yaslioglu, M. Murat
Publié: (2025)
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
par: Javadi, Farnoosh, et autres
Publié: (2023)
par: Javadi, Farnoosh, et autres
Publié: (2023)
Element-wise Attention Is All You Need
par: Feng, Guoxin
Publié: (2025)
par: Feng, Guoxin
Publié: (2025)
Exploring the Integration of Key-Value Attention Into Pure and Hybrid Transformers for Semantic Segmentation
par: Hwa, DeShin, et autres
Publié: (2025)
par: Hwa, DeShin, et autres
Publié: (2025)
Thin Keys, Full Values: Reducing KV Cache via Low-Dimensional Attention Selection
par: Yao, Hengshuai, et autres
Publié: (2026)
par: Yao, Hengshuai, et autres
Publié: (2026)
Tensor Product Attention Is All You Need
par: Zhang, Yifan, et autres
Publié: (2025)
par: Zhang, Yifan, et autres
Publié: (2025)
Attention Smoothing Is All You Need For Unlearning
par: Zade, Saleh Zare, et autres
Publié: (2026)
par: Zade, Saleh Zare, et autres
Publié: (2026)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
par: Yue, Yuxuan, et autres
Publié: (2024)
par: Yue, Yuxuan, et autres
Publié: (2024)
Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
par: Wi, Hyowon, et autres
Publié: (2025)
par: Wi, Hyowon, et autres
Publié: (2025)
Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers
par: Horton, Mark, et autres
Publié: (2025)
par: Horton, Mark, et autres
Publié: (2025)
RecurFormer: Not All Transformer Heads Need Self-Attention
par: Yan, Ruiqing, et autres
Publié: (2024)
par: Yan, Ruiqing, et autres
Publié: (2024)
TransMLA: Multi-Head Latent Attention Is All You Need
par: Meng, Fanxu, et autres
Publié: (2025)
par: Meng, Fanxu, et autres
Publié: (2025)
Context is All You Need
par: Delanois, Jean Erik, et autres
Publié: (2026)
par: Delanois, Jean Erik, et autres
Publié: (2026)
LOOKAT: Lookup-Optimized Key-Attention for Memory-Efficient Transformers
par: Karmore, Aryan
Publié: (2026)
par: Karmore, Aryan
Publié: (2026)
Attention Is All You Need for KV Cache in Diffusion LLMs
par: Nguyen-Tri, Quan, et autres
Publié: (2025)
par: Nguyen-Tri, Quan, et autres
Publié: (2025)
Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
par: Gao, Wei, et autres
Publié: (2025)
par: Gao, Wei, et autres
Publié: (2025)
Exploitation Is All You Need... for Exploration
par: Rentschler, Micah, et autres
Publié: (2025)
par: Rentschler, Micah, et autres
Publié: (2025)
The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference
par: Qasim, Kaleem Ullah, et autres
Publié: (2026)
par: Qasim, Kaleem Ullah, et autres
Publié: (2026)
What Matters in Transformers? Not All Attention is Needed
par: He, Shwai, et autres
Publié: (2024)
par: He, Shwai, et autres
Publié: (2024)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
par: Yang, Xu, et autres
Publié: (2026)
par: Yang, Xu, et autres
Publié: (2026)
All You Need Is Synthetic Task Augmentation
par: Godin, Guillaume
Publié: (2025)
par: Godin, Guillaume
Publié: (2025)
VSFormer: Value and Shape-Aware Transformer with Prior-Enhanced Self-Attention for Multivariate Time Series Classification
par: Xi, Wenjie, et autres
Publié: (2024)
par: Xi, Wenjie, et autres
Publié: (2024)
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
par: Bu, Rui, et autres
Publié: (2025)
par: Bu, Rui, et autres
Publié: (2025)
Mamba or Transformer for Time Series Forecasting? Mixture of Universals (MoU) Is All You Need
par: Peng, Sijia, et autres
Publié: (2024)
par: Peng, Sijia, et autres
Publié: (2024)
Transduction is All You Need for Structured Data Workflows
par: Gliozzo, Alfio, et autres
Publié: (2025)
par: Gliozzo, Alfio, et autres
Publié: (2025)
Holographic Transformers for Complex-Valued Signal Processing: Integrating Phase Interference into Self-Attention
par: Huang, Enhao, et autres
Publié: (2025)
par: Huang, Enhao, et autres
Publié: (2025)
Attention Is Not What You Need
par: Chong, Zhang
Publié: (2025)
par: Chong, Zhang
Publié: (2025)
Multi-objective Optimization in CPU Design Space Exploration: Attention is All You Need
par: Xue, Runzhen, et autres
Publié: (2024)
par: Xue, Runzhen, et autres
Publié: (2024)
More Agents Is All You Need
par: Li, Junyou, et autres
Publié: (2024)
par: Li, Junyou, et autres
Publié: (2024)
Cooperation Is All You Need
par: Adeel, Ahsan, et autres
Publié: (2023)
par: Adeel, Ahsan, et autres
Publié: (2023)
Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning
par: Yu, Zhiyin, et autres
Publié: (2026)
par: Yu, Zhiyin, et autres
Publié: (2026)
Capabilities Ain't All You Need: Measuring Propensities in AI
par: Romero-Alvarado, Daniel, et autres
Publié: (2026)
par: Romero-Alvarado, Daniel, et autres
Publié: (2026)
HDL-GPT: High-Quality HDL is All You Need
par: Kumar, Bhuvnesh, et autres
Publié: (2024)
par: Kumar, Bhuvnesh, et autres
Publié: (2024)
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
par: Goldstein, Daniel, et autres
Publié: (2026)
par: Goldstein, Daniel, et autres
Publié: (2026)
ZoomR: Memory Efficient Reasoning through Multi-Granularity Key Value Retrieval
par: Yang, David H., et autres
Publié: (2026)
par: Yang, David H., et autres
Publié: (2026)
Is Diversity All You Need for Scalable Robotic Manipulation?
par: Shi, Modi, et autres
Publié: (2025)
par: Shi, Modi, et autres
Publié: (2025)
Confidence Is All You Need for MI Attacks
par: Sinha, Abhishek, et autres
Publié: (2023)
par: Sinha, Abhishek, et autres
Publié: (2023)
Context-Selective State Space Models: Feedback is All You Need
par: Zattra, Riccardo, et autres
Publié: (2025)
par: Zattra, Riccardo, et autres
Publié: (2025)
Documents similaires
-
Can an MLP Absorb Its Own Skip Connection?
par: Mijoski, Antonij, et autres
Publié: (2026) -
Beyond Linearity in Attention Projections: The Case for Nonlinear Queries
par: Karbevski, Marko
Publié: (2026) -
Attention is All You Need Until You Need Retention
par: Yaslioglu, M. Murat
Publié: (2025) -
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
par: Javadi, Farnoosh, et autres
Publié: (2023) -
Element-wise Attention Is All You Need
par: Feng, Guoxin
Publié: (2025)