Coupled Query-Key Dynamics for Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gahtan, Barak, Bronstein, Alex M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Effective Sample Size and Generalization Bounds for Temporal Networks
von: Gahtan, Barak, et al.
Veröffentlicht: (2025)
von: Gahtan, Barak, et al.
Veröffentlicht: (2025)
Exploring QUIC Dynamics: A Large-Scale Dataset for Encrypted Traffic Analysis
von: Gahtan, Barak, et al.
Veröffentlicht: (2024)
von: Gahtan, Barak, et al.
Veröffentlicht: (2024)
Data-Driven Cellular Network Selector for Vehicle Teleoperations
von: Gahtan, Barak, et al.
Veröffentlicht: (2024)
von: Gahtan, Barak, et al.
Veröffentlicht: (2024)
Estimating the Number of HTTP/3 Responses in QUIC Using Deep Learning
von: Gahtan, Barak, et al.
Veröffentlicht: (2024)
von: Gahtan, Barak, et al.
Veröffentlicht: (2024)
From Lab to Wrist: Bridging Metabolic Monitoring and Consumer Wearables for Heart Rate and Oxygen Consumption Modeling
von: Gahtan, Barak, et al.
Veröffentlicht: (2025)
von: Gahtan, Barak, et al.
Veröffentlicht: (2025)
WearableMil: An End-to-End Framework for Military Activity Recognition and Performance Monitoring
von: Gahtan, Barak, et al.
Veröffentlicht: (2024)
von: Gahtan, Barak, et al.
Veröffentlicht: (2024)
Beyond the Alphabet: Deep Signal Embedding for Enhanced DNA Clustering
von: Abraham, Hadas, et al.
Veröffentlicht: (2024)
von: Abraham, Hadas, et al.
Veröffentlicht: (2024)
Wildfire Simulation with Differentiable Randers-Finsler Eikonal Solvers
von: Gahtan, Barak, et al.
Veröffentlicht: (2026)
von: Gahtan, Barak, et al.
Veröffentlicht: (2026)
Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
von: Filipek, Adam
Veröffentlicht: (2025)
von: Filipek, Adam
Veröffentlicht: (2025)
Causal Attention with Lookahead Keys
von: Song, Zhuoqing, et al.
Veröffentlicht: (2025)
von: Song, Zhuoqing, et al.
Veröffentlicht: (2025)
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
von: Le, Dong, et al.
Veröffentlicht: (2026)
von: Le, Dong, et al.
Veröffentlicht: (2026)
Trellis: Learning to Compress Key-Value Memory in Attention Models
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
von: Brandon, William, et al.
Veröffentlicht: (2024)
von: Brandon, William, et al.
Veröffentlicht: (2024)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric Learning
von: Yu, Zeping, et al.
Veröffentlicht: (2024)
von: Yu, Zeping, et al.
Veröffentlicht: (2024)
Quantifying Logical Consistency in Transformers via Query-Key Alignment
von: Tulchinskii, Eduard, et al.
Veröffentlicht: (2025)
von: Tulchinskii, Eduard, et al.
Veröffentlicht: (2025)
Decomposing Attention To Find Context-Sensitive Neurons
von: Gibson, Alex
Veröffentlicht: (2025)
von: Gibson, Alex
Veröffentlicht: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
von: Xiao, Da, et al.
Veröffentlicht: (2024)
von: Xiao, Da, et al.
Veröffentlicht: (2024)
LLMs can hide text in other text of the same length
von: Norelli, Antonio, et al.
Veröffentlicht: (2025)
von: Norelli, Antonio, et al.
Veröffentlicht: (2025)
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
von: Javadi, Farnoosh, et al.
Veröffentlicht: (2023)
von: Javadi, Farnoosh, et al.
Veröffentlicht: (2023)
Adaptive Querying with AI Persona Priors
von: Wang, Kaizheng, et al.
Veröffentlicht: (2026)
von: Wang, Kaizheng, et al.
Veröffentlicht: (2026)
Inducing Meaningful Units from Character Sequences with Dynamic Capacity Slot Attention
von: Behjati, Melika, et al.
Veröffentlicht: (2021)
von: Behjati, Melika, et al.
Veröffentlicht: (2021)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
Dynamic Multimodal Sentiment Analysis: Leveraging Cross-Modal Attention for Enabled Classification
von: Lee, Hui, et al.
Veröffentlicht: (2025)
von: Lee, Hui, et al.
Veröffentlicht: (2025)
On the Query Complexity of Verifier-Assisted Language Generation
von: Botta, Edoardo, et al.
Veröffentlicht: (2025)
von: Botta, Edoardo, et al.
Veröffentlicht: (2025)
Query-Focused Extractive Summarization for Sentiment Explanation
von: Moubtahij, Ahmed, et al.
Veröffentlicht: (2025)
von: Moubtahij, Ahmed, et al.
Veröffentlicht: (2025)
Trainable Dynamic Mask Sparse Attention
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
von: Song, Guanghui, et al.
Veröffentlicht: (2025)
von: Song, Guanghui, et al.
Veröffentlicht: (2025)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
von: Huang, Yanwen, et al.
Veröffentlicht: (2025)
von: Huang, Yanwen, et al.
Veröffentlicht: (2025)
Why Softmax Attention Outperforms Linear Attention
von: Deng, Yichuan, et al.
Veröffentlicht: (2023)
von: Deng, Yichuan, et al.
Veröffentlicht: (2023)
Beyond Uniform Query Distribution: Key-Driven Grouped Query Attention
von: Khan, Zohaib, et al.
Veröffentlicht: (2024)
von: Khan, Zohaib, et al.
Veröffentlicht: (2024)
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2024)
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2024)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
SEA: Sparse Linear Attention with Estimated Attention Mask
von: Lee, Heejun, et al.
Veröffentlicht: (2023)
von: Lee, Heejun, et al.
Veröffentlicht: (2023)
Are Large Language Models Good Temporal Graph Learners?
von: Huang, Shenyang, et al.
Veröffentlicht: (2025)
von: Huang, Shenyang, et al.
Veröffentlicht: (2025)
Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
von: Qin, Guanghui, et al.
Veröffentlicht: (2021)
von: Qin, Guanghui, et al.
Veröffentlicht: (2021)
QueryBuilder: Human-in-the-Loop Query Development for Information Retrieval
von: Kandula, Hemanth, et al.
Veröffentlicht: (2024)
von: Kandula, Hemanth, et al.
Veröffentlicht: (2024)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
von: Wong, Liang Ze
Veröffentlicht: (2025)
von: Wong, Liang Ze
Veröffentlicht: (2025)
Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
von: Erden, Caner
Veröffentlicht: (2025)
von: Erden, Caner
Veröffentlicht: (2025)
Ähnliche Einträge
-
Effective Sample Size and Generalization Bounds for Temporal Networks
von: Gahtan, Barak, et al.
Veröffentlicht: (2025) -
Exploring QUIC Dynamics: A Large-Scale Dataset for Encrypted Traffic Analysis
von: Gahtan, Barak, et al.
Veröffentlicht: (2024) -
Data-Driven Cellular Network Selector for Vehicle Teleoperations
von: Gahtan, Barak, et al.
Veröffentlicht: (2024) -
Estimating the Number of HTTP/3 Responses in QUIC Using Deep Learning
von: Gahtan, Barak, et al.
Veröffentlicht: (2024) -
From Lab to Wrist: Bridging Metabolic Monitoring and Consumer Wearables for Heart Rate and Oxygen Consumption Modeling
von: Gahtan, Barak, et al.
Veröffentlicht: (2025)