Coupled Query-Key Dynamics for Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Gahtan, Barak, Bronstein, Alex M. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Effective Sample Size and Generalization Bounds for Temporal Networks
by: Gahtan, Barak, et al.
Published: (2025)
by: Gahtan, Barak, et al.
Published: (2025)
Exploring QUIC Dynamics: A Large-Scale Dataset for Encrypted Traffic Analysis
by: Gahtan, Barak, et al.
Published: (2024)
by: Gahtan, Barak, et al.
Published: (2024)
Data-Driven Cellular Network Selector for Vehicle Teleoperations
by: Gahtan, Barak, et al.
Published: (2024)
by: Gahtan, Barak, et al.
Published: (2024)
Estimating the Number of HTTP/3 Responses in QUIC Using Deep Learning
by: Gahtan, Barak, et al.
Published: (2024)
by: Gahtan, Barak, et al.
Published: (2024)
From Lab to Wrist: Bridging Metabolic Monitoring and Consumer Wearables for Heart Rate and Oxygen Consumption Modeling
by: Gahtan, Barak, et al.
Published: (2025)
by: Gahtan, Barak, et al.
Published: (2025)
WearableMil: An End-to-End Framework for Military Activity Recognition and Performance Monitoring
by: Gahtan, Barak, et al.
Published: (2024)
by: Gahtan, Barak, et al.
Published: (2024)
Beyond the Alphabet: Deep Signal Embedding for Enhanced DNA Clustering
by: Abraham, Hadas, et al.
Published: (2024)
by: Abraham, Hadas, et al.
Published: (2024)
Wildfire Simulation with Differentiable Randers-Finsler Eikonal Solvers
by: Gahtan, Barak, et al.
Published: (2026)
by: Gahtan, Barak, et al.
Published: (2026)
Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
by: Filipek, Adam
Published: (2025)
by: Filipek, Adam
Published: (2025)
Causal Attention with Lookahead Keys
by: Song, Zhuoqing, et al.
Published: (2025)
by: Song, Zhuoqing, et al.
Published: (2025)
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
by: Le, Dong, et al.
Published: (2026)
by: Le, Dong, et al.
Published: (2026)
Trellis: Learning to Compress Key-Value Memory in Attention Models
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
by: Brandon, William, et al.
Published: (2024)
by: Brandon, William, et al.
Published: (2024)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
by: Chen, Yingfa, et al.
Published: (2025)
by: Chen, Yingfa, et al.
Published: (2025)
How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric Learning
by: Yu, Zeping, et al.
Published: (2024)
by: Yu, Zeping, et al.
Published: (2024)
Quantifying Logical Consistency in Transformers via Query-Key Alignment
by: Tulchinskii, Eduard, et al.
Published: (2025)
by: Tulchinskii, Eduard, et al.
Published: (2025)
Decomposing Attention To Find Context-Sensitive Neurons
by: Gibson, Alex
Published: (2025)
by: Gibson, Alex
Published: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
by: Xiao, Da, et al.
Published: (2024)
by: Xiao, Da, et al.
Published: (2024)
LLMs can hide text in other text of the same length
by: Norelli, Antonio, et al.
Published: (2025)
by: Norelli, Antonio, et al.
Published: (2025)
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
by: Javadi, Farnoosh, et al.
Published: (2023)
by: Javadi, Farnoosh, et al.
Published: (2023)
Adaptive Querying with AI Persona Priors
by: Wang, Kaizheng, et al.
Published: (2026)
by: Wang, Kaizheng, et al.
Published: (2026)
Inducing Meaningful Units from Character Sequences with Dynamic Capacity Slot Attention
by: Behjati, Melika, et al.
Published: (2021)
by: Behjati, Melika, et al.
Published: (2021)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
by: Zarch, Hossein Entezari, et al.
Published: (2025)
by: Zarch, Hossein Entezari, et al.
Published: (2025)
Dynamic Multimodal Sentiment Analysis: Leveraging Cross-Modal Attention for Enabled Classification
by: Lee, Hui, et al.
Published: (2025)
by: Lee, Hui, et al.
Published: (2025)
On the Query Complexity of Verifier-Assisted Language Generation
by: Botta, Edoardo, et al.
Published: (2025)
by: Botta, Edoardo, et al.
Published: (2025)
Query-Focused Extractive Summarization for Sentiment Explanation
by: Moubtahij, Ahmed, et al.
Published: (2025)
by: Moubtahij, Ahmed, et al.
Published: (2025)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
by: Song, Guanghui, et al.
Published: (2025)
by: Song, Guanghui, et al.
Published: (2025)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
by: Huang, Yanwen, et al.
Published: (2025)
by: Huang, Yanwen, et al.
Published: (2025)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
Beyond Uniform Query Distribution: Key-Driven Grouped Query Attention
by: Khan, Zohaib, et al.
Published: (2024)
by: Khan, Zohaib, et al.
Published: (2024)
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
by: Jiang, Huiqiang, et al.
Published: (2024)
by: Jiang, Huiqiang, et al.
Published: (2024)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
by: Tang, Hongyin, et al.
Published: (2024)
by: Tang, Hongyin, et al.
Published: (2024)
SEA: Sparse Linear Attention with Estimated Attention Mask
by: Lee, Heejun, et al.
Published: (2023)
by: Lee, Heejun, et al.
Published: (2023)
Are Large Language Models Good Temporal Graph Learners?
by: Huang, Shenyang, et al.
Published: (2025)
by: Huang, Shenyang, et al.
Published: (2025)
Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
by: Qin, Guanghui, et al.
Published: (2021)
by: Qin, Guanghui, et al.
Published: (2021)
QueryBuilder: Human-in-the-Loop Query Development for Information Retrieval
by: Kandula, Hemanth, et al.
Published: (2024)
by: Kandula, Hemanth, et al.
Published: (2024)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
by: Ildiz, M. Emrullah, et al.
Published: (2024)
by: Ildiz, M. Emrullah, et al.
Published: (2024)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
by: Wong, Liang Ze
Published: (2025)
by: Wong, Liang Ze
Published: (2025)
Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
by: Erden, Caner
Published: (2025)
by: Erden, Caner
Published: (2025)
Similar Items
-
Effective Sample Size and Generalization Bounds for Temporal Networks
by: Gahtan, Barak, et al.
Published: (2025) -
Exploring QUIC Dynamics: A Large-Scale Dataset for Encrypted Traffic Analysis
by: Gahtan, Barak, et al.
Published: (2024) -
Data-Driven Cellular Network Selector for Vehicle Teleoperations
by: Gahtan, Barak, et al.
Published: (2024) -
Estimating the Number of HTTP/3 Responses in QUIC Using Deep Learning
by: Gahtan, Barak, et al.
Published: (2024) -
From Lab to Wrist: Bridging Metabolic Monitoring and Consumer Wearables for Heart Rate and Oxygen Consumption Modeling
by: Gahtan, Barak, et al.
Published: (2025)