Weighted Grouped Query Attention in Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chinnakonduru, Sai Sena, Mohapatra, Astarag |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
QCQA: Quality and Capacity-aware grouped Query Attention
von: Joshi, Vinay, et al.
Veröffentlicht: (2024)
von: Joshi, Vinay, et al.
Veröffentlicht: (2024)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
GTA: Grouped-head latenT Attention
von: Sun, Luoyang, et al.
Veröffentlicht: (2025)
von: Sun, Luoyang, et al.
Veröffentlicht: (2025)
Transformers for Complex Query Answering over Knowledge Hypergraphs
von: Tsang, Hong Ting, et al.
Veröffentlicht: (2025)
von: Tsang, Hong Ting, et al.
Veröffentlicht: (2025)
An Investigation on Group Query Hallucination Attacks
von: Miao, Kehao, et al.
Veröffentlicht: (2025)
von: Miao, Kehao, et al.
Veröffentlicht: (2025)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
von: Zhu, Shenzhe
Veröffentlicht: (2025)
von: Zhu, Shenzhe
Veröffentlicht: (2025)
Memorization in Attention-only Transformers
von: Dana, Léo, et al.
Veröffentlicht: (2024)
von: Dana, Léo, et al.
Veröffentlicht: (2024)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2025)
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2025)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
von: Bae, Jeongin, et al.
Veröffentlicht: (2026)
von: Bae, Jeongin, et al.
Veröffentlicht: (2026)
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
von: Javadi, Farnoosh, et al.
Veröffentlicht: (2023)
von: Javadi, Farnoosh, et al.
Veröffentlicht: (2023)
PGB: One-Shot Pruning for BERT via Weight Grouping and Permutation
von: Lim, Hyemin, et al.
Veröffentlicht: (2025)
von: Lim, Hyemin, et al.
Veröffentlicht: (2025)
Hierarchical vs. Flat Iteration in Shared-Weight Transformers
von: Han, Sang-Il
Veröffentlicht: (2026)
von: Han, Sang-Il
Veröffentlicht: (2026)
SQL-Exchange: Transforming SQL Queries Across Domains
von: Daviran, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Daviran, Mohammadreza, et al.
Veröffentlicht: (2025)
More Expressive Attention with Negative Weights
von: Lv, Ang, et al.
Veröffentlicht: (2024)
von: Lv, Ang, et al.
Veröffentlicht: (2024)
Enhancing Essay Scoring with Adversarial Weights Perturbation and Metric-specific AttentionPooling
von: Huang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Huang, Jiaxin, et al.
Veröffentlicht: (2024)
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias
von: Chhabra, Anshuman, et al.
Veröffentlicht: (2024)
von: Chhabra, Anshuman, et al.
Veröffentlicht: (2024)
QueryNER: Segmentation of E-commerce Queries
von: Palen-Michel, Chester, et al.
Veröffentlicht: (2024)
von: Palen-Michel, Chester, et al.
Veröffentlicht: (2024)
The Attentional White Bear Effect in Transformer Language Models
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2026)
von: Ramnauth, Rebecca, et al.
Veröffentlicht: (2026)
Crisp Attention: Regularizing Transformers via Structured Sparsity
von: Gandhi, Sagar, et al.
Veröffentlicht: (2025)
von: Gandhi, Sagar, et al.
Veröffentlicht: (2025)
Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection
von: Miralles-González, Pablo, et al.
Veröffentlicht: (2025)
von: Miralles-González, Pablo, et al.
Veröffentlicht: (2025)
Simulating Weighted Automata over Sequences and Trees with Transformers
von: Rizvi, Michael, et al.
Veröffentlicht: (2024)
von: Rizvi, Michael, et al.
Veröffentlicht: (2024)
SignAttention: On the Interpretability of Transformer Models for Sign Language Translation
von: Bianco, Pedro Alejandro Dal, et al.
Veröffentlicht: (2024)
von: Bianco, Pedro Alejandro Dal, et al.
Veröffentlicht: (2024)
No Query, No Access
von: Wang, Wenqiang, et al.
Veröffentlicht: (2025)
von: Wang, Wenqiang, et al.
Veröffentlicht: (2025)
Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation
von: Liu, Zekang, et al.
Veröffentlicht: (2025)
von: Liu, Zekang, et al.
Veröffentlicht: (2025)
Selective Attention Improves Transformer
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2024)
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2024)
Generating Query-Focused Summarization Datasets from Query-Free Summarization Datasets
von: Chali, Yllias, et al.
Veröffentlicht: (2026)
von: Chali, Yllias, et al.
Veröffentlicht: (2026)
Mass-Editing Memory with Attention in Transformers: A cross-lingual exploration of knowledge
von: Tamayo, Daniel, et al.
Veröffentlicht: (2025)
von: Tamayo, Daniel, et al.
Veröffentlicht: (2025)
ShishuLM : Achieving Optimal and Efficient Parameterization with Low Attention Transformer Models
von: Kumar, Shivanshu, et al.
Veröffentlicht: (2025)
von: Kumar, Shivanshu, et al.
Veröffentlicht: (2025)
Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models
von: Tian, Yuxing, et al.
Veröffentlicht: (2026)
von: Tian, Yuxing, et al.
Veröffentlicht: (2026)
Uncovering the Role of Initial Saliency in U-Shaped Attention Bias: Scaling Initial Token Weight for Enhanced Long-Text Processing
von: Qiang, Zewen, et al.
Veröffentlicht: (2025)
von: Qiang, Zewen, et al.
Veröffentlicht: (2025)
QueryPlot: Generating Geological Evidence Layers using Natural Language Queries for Mineral Exploration
von: Ye, Meng, et al.
Veröffentlicht: (2026)
von: Ye, Meng, et al.
Veröffentlicht: (2026)
Query-Efficient Planning with Language Models
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2024)
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2024)
What Matters in Transformers? Not All Attention is Needed
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
Arabic Tweet Act: A Weighted Ensemble Pre-Trained Transformer Model for Classifying Arabic Speech Acts on Twitter
von: Alshehri, Khadejaa, et al.
Veröffentlicht: (2024)
von: Alshehri, Khadejaa, et al.
Veröffentlicht: (2024)
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
von: Ji, Tao, et al.
Veröffentlicht: (2025)
von: Ji, Tao, et al.
Veröffentlicht: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
von: Balogh, Peter
Veröffentlicht: (2026)
von: Balogh, Peter
Veröffentlicht: (2026)
Ähnliche Einträge
-
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
von: Chen, Yingfa, et al.
Veröffentlicht: (2025) -
QCQA: Quality and Capacity-aware grouped Query Attention
von: Joshi, Vinay, et al.
Veröffentlicht: (2024) -
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
von: Devoto, Alessio, et al.
Veröffentlicht: (2025) -
GTA: Grouped-head latenT Attention
von: Sun, Luoyang, et al.
Veröffentlicht: (2025) -
Transformers for Complex Query Answering over Knowledge Hypergraphs
von: Tsang, Hong Ting, et al.
Veröffentlicht: (2025)