GQA-μP: The maximal parameterization update for grouped query attention
Fuente:
arXiv
Saved in:
| Main Authors: | Chickering, Kyle R., Wang, Huijuan, Wu, Mengxi, Moreno, Alexander, Chen, Muhao, Ma, Xuezhe, Soboleva, Daria, Hestness, Joel, Liu, Zhengzhong, Xing, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparse maximal update parameterization: A holistic approach to sparse training dynamics
by: Dey, Nolan, et al.
Published: (2024)
by: Dey, Nolan, et al.
Published: (2024)
QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining
by: Chickering, Kyle R., et al.
Published: (2025)
by: Chickering, Kyle R., et al.
Published: (2025)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
SlimPajama-DC: Understanding Data Combinations for LLM Training
by: Shen, Zhiqiang, et al.
Published: (2023)
by: Shen, Zhiqiang, et al.
Published: (2023)
μP$^2$: Effective Sharpness Aware Minimization Requires Layerwise Perturbation Scaling
by: Haas, Moritz, et al.
Published: (2024)
by: Haas, Moritz, et al.
Published: (2024)
A Quasilinear Algorithm for Computing Higher-Order Derivatives of Deep Feed-Forward Neural Networks
by: Chickering, Kyle R.
Published: (2024)
by: Chickering, Kyle R.
Published: (2024)
EMO: Frustratingly Easy Progressive Training of Extendable MoE
by: Jin, Linghao, et al.
Published: (2026)
by: Jin, Linghao, et al.
Published: (2026)
Agent-Based Insight into Eco-Choices: Simulating the Fast Fashion Shift
by: Soboleva, Daria, et al.
Published: (2024)
by: Soboleva, Daria, et al.
Published: (2024)
Asymptotically self-similar shock formation for 1d fractal Burgers equation
by: Chickering, Kyle R., et al.
Published: (2021)
by: Chickering, Kyle R., et al.
Published: (2021)
Library Service to Spanish Speaking Patrons: A Practical Guide.
by: Moller, Sharon Chickering
Published: (2001)
by: Moller, Sharon Chickering
Published: (2001)
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Sensitivity-Positional Co-Localization in GQA Transformers
by: Rao, Manoj Chandrashekar
Published: (2026)
by: Rao, Manoj Chandrashekar
Published: (2026)
QUEST: A robust attention formulation using query-modulated spherical attention
by: Govindarajan, Hariprasath, et al.
Published: (2026)
by: Govindarajan, Hariprasath, et al.
Published: (2026)
Asymmetric Idiosyncrasies in Multimodal Models
by: Tao, Muzi, et al.
Published: (2026)
by: Tao, Muzi, et al.
Published: (2026)
Enumeration and updates for conjunctive linear algebra queries through expressibility
by: Muñoz, Thomas, et al.
Published: (2023)
by: Muñoz, Thomas, et al.
Published: (2023)
Crystal: Illuminating LLM Abilities on Language and Code
by: Tao, Tianhua, et al.
Published: (2024)
by: Tao, Tianhua, et al.
Published: (2024)
Lower bounds for graph reconstruction with maximal independent set queries
by: Michel, Lukas, et al.
Published: (2024)
by: Michel, Lukas, et al.
Published: (2024)
The Balancing Exercise and the Resolution of Rhetorical Antinomies in Judicial Decision‐Making
by: Anita Soboleva
Published: (2024)
by: Anita Soboleva
Published: (2024)
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
by: Gray, Gavia, et al.
Published: (2024)
by: Gray, Gavia, et al.
Published: (2024)
Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
by: Zhu, Tinghui, et al.
Published: (2024)
by: Zhu, Tinghui, et al.
Published: (2024)
Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA
by: Gumaan, Esmail
Published: (2025)
by: Gumaan, Esmail
Published: (2025)
PatentEdits: Framing Patent Novelty as Textual Entailment
by: Lee, Ryan, et al.
Published: (2024)
by: Lee, Ryan, et al.
Published: (2024)
Nuevo sistema de incentivos de la Unión Soviética: estudio de su funcionamiento en Kief
by: Galina Danilovna Soboleva
Published: (1970)
by: Galina Danilovna Soboleva
Published: (1970)
The new soviet incentive system: a study of its operation in Kiev
by: Galina Danilovna Soboleva
Published: (1970)
by: Galina Danilovna Soboleva
Published: (1970)
Plasma recombination in runaway discharges in tokamak TCABR
by: T.K. Soboleva
Published: (2002)
by: T.K. Soboleva
Published: (2002)
Le nouveau système soviétique de stimulation: etude de son fonctionnement à Kiev
by: Galina Danilovna Soboleva
Published: (1970)
by: Galina Danilovna Soboleva
Published: (1970)
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA
by: Jin, Qingyun, et al.
Published: (2024)
by: Jin, Qingyun, et al.
Published: (2024)
Reclassification queries in a geographical data warehouse
by: Francisco J. Moreno
Published: (2010)
by: Francisco J. Moreno
Published: (2010)
K2-V2: A 360-Open, Reasoning-Enhanced LLM
by: K2 Team, et al.
Published: (2025)
by: K2 Team, et al.
Published: (2025)
Efficiently Disentangle Causal Representations
by: Li, Yuanpeng, et al.
Published: (2022)
by: Li, Yuanpeng, et al.
Published: (2022)
On the orbits of a finite solvable primitive linear group
by: Yang, Yong, et al.
Published: (2024)
by: Yang, Yong, et al.
Published: (2024)
Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities
by: Wen, Nuan, et al.
Published: (2026)
by: Wen, Nuan, et al.
Published: (2026)
LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems
by: Xu, Nan, et al.
Published: (2024)
by: Xu, Nan, et al.
Published: (2024)
DecoPrompt : Decoding Prompts Reduces Hallucinations when Large Language Models Meet False Premises
by: Xu, Nan, et al.
Published: (2024)
by: Xu, Nan, et al.
Published: (2024)
From Orientalism to Cultural Capital
by: Soboleva, Olga, et al.
Published: (2017)
by: Soboleva, Olga, et al.
Published: (2017)
Probing maximal flavor changing $Z'$ in $U(1)_{L_μ-L_τ}$ at $μ$TRISTAN
by: Huang, Fei, et al.
Published: (2025)
by: Huang, Fei, et al.
Published: (2025)
One other parameterization of SU(4) group
by: Khvedelidze, Arsen, et al.
Published: (2024)
by: Khvedelidze, Arsen, et al.
Published: (2024)
Robust forecast aggregation via additional queries
by: Frongillo, Rafael, et al.
Published: (2025)
by: Frongillo, Rafael, et al.
Published: (2025)
Similar Items
-
Sparse maximal update parameterization: A holistic approach to sparse training dynamics
by: Dey, Nolan, et al.
Published: (2024) -
QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining
by: Chickering, Kyle R., et al.
Published: (2025) -
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
by: Bergsma, Shane, et al.
Published: (2025) -
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
by: Bergsma, Shane, et al.
Published: (2025) -
SlimPajama-DC: Understanding Data Combinations for LLM Training
by: Shen, Zhiqiang, et al.
Published: (2023)