Scale-invariant Attention
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Anson, Ben, Wang, Xi, Aitchison, Laurence |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Controlling changes to attention logits
par: Anson, Ben, et autres
Publié: (2025)
par: Anson, Ben, et autres
Publié: (2025)
Learning to Skip the Middle Layers of Transformers
par: Lawson, Tim, et autres
Publié: (2025)
par: Lawson, Tim, et autres
Publié: (2025)
Batch size invariant Adam
par: Wang, Xi, et autres
Publié: (2024)
par: Wang, Xi, et autres
Publié: (2024)
Function-Space Learning Rates
par: Milsom, Edward, et autres
Publié: (2025)
par: Milsom, Edward, et autres
Publié: (2025)
Flexible Infinite-Width Graph Convolutional Neural Networks
par: Anson, Ben, et autres
Publié: (2024)
par: Anson, Ben, et autres
Publié: (2024)
Convolutional Deep Kernel Machines
par: Milsom, Edward, et autres
Publié: (2023)
par: Milsom, Edward, et autres
Publié: (2023)
Stochastic Kernel Regularisation Improves Generalisation in Deep Kernel Machines
par: Milsom, Edward, et autres
Publié: (2024)
par: Milsom, Edward, et autres
Publié: (2024)
Residual Stream Analysis with Multi-Layer SAEs
par: Lawson, Tim, et autres
Publié: (2024)
par: Lawson, Tim, et autres
Publié: (2024)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
par: Farnik, Lucy, et autres
Publié: (2025)
par: Farnik, Lucy, et autres
Publié: (2025)
Questionable practices in machine learning
par: Leech, Gavin, et autres
Publié: (2024)
par: Leech, Gavin, et autres
Publié: (2024)
How to set AdamW's weight decay as you scale model and dataset size
par: Wang, Xi, et autres
Publié: (2024)
par: Wang, Xi, et autres
Publié: (2024)
Why you don't overfit, and don't need Bayes if you only train for one epoch
par: Aitchison, Laurence
Publié: (2024)
par: Aitchison, Laurence
Publié: (2024)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
par: Chen, Lida, et autres
Publié: (2025)
par: Chen, Lida, et autres
Publié: (2025)
Scaling Linear Attention with Sparse State Expansion
par: Pan, Yuqi, et autres
Publié: (2025)
par: Pan, Yuqi, et autres
Publié: (2025)
Scaling Reasoning without Attention
par: Zhao, Xueliang, et autres
Publié: (2025)
par: Zhao, Xueliang, et autres
Publié: (2025)
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
par: MiniMax, et autres
Publié: (2025)
par: MiniMax, et autres
Publié: (2025)
Learnable Multi-Scale Wavelet Transformer: A Novel Alternative to Self-Attention
par: Kiruluta, Andrew, et autres
Publié: (2025)
par: Kiruluta, Andrew, et autres
Publié: (2025)
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
par: Wiegand, Götz-Henrik, et autres
Publié: (2026)
par: Wiegand, Götz-Henrik, et autres
Publié: (2026)
Learning to Attribute with Attention
par: Cohen-Wang, Benjamin, et autres
Publié: (2025)
par: Cohen-Wang, Benjamin, et autres
Publié: (2025)
Stacking Small Language Models for Generalizability
par: Liang, Laurence
Publié: (2024)
par: Liang, Laurence
Publié: (2024)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
par: Qiu, Quantong, et autres
Publié: (2026)
par: Qiu, Quantong, et autres
Publié: (2026)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
par: Ma, Qingsen, et autres
Publié: (2026)
par: Ma, Qingsen, et autres
Publié: (2026)
Scaling Bidirectional Spans and Span Violations in Attention Mechanism
par: Kim, Jongwook, et autres
Publié: (2025)
par: Kim, Jongwook, et autres
Publié: (2025)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
par: Lee, Changhun, et autres
Publié: (2025)
par: Lee, Changhun, et autres
Publié: (2025)
Bayesian Low-rank Adaptation for Large Language Models
par: Yang, Adam X., et autres
Publié: (2023)
par: Yang, Adam X., et autres
Publié: (2023)
LoMA: Lossless Compressed Memory Attention
par: Wang, Yumeng, et autres
Publié: (2024)
par: Wang, Yumeng, et autres
Publié: (2024)
Why Softmax Attention Outperforms Linear Attention
par: Deng, Yichuan, et autres
Publié: (2023)
par: Deng, Yichuan, et autres
Publié: (2023)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
par: Dentamaro, Vincenzo
Publié: (2025)
par: Dentamaro, Vincenzo
Publié: (2025)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
par: Song, Dinghong, et autres
Publié: (2025)
par: Song, Dinghong, et autres
Publié: (2025)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
par: Fu, Zizhuo, et autres
Publié: (2026)
par: Fu, Zizhuo, et autres
Publié: (2026)
Multi-matrix Factorization Attention
par: Hu, Jingcheng, et autres
Publié: (2024)
par: Hu, Jingcheng, et autres
Publié: (2024)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
par: Tan, Shawn, et autres
Publié: (2024)
par: Tan, Shawn, et autres
Publié: (2024)
SEA: Sparse Linear Attention with Estimated Attention Mask
par: Lee, Heejun, et autres
Publié: (2023)
par: Lee, Heejun, et autres
Publié: (2023)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
par: Xi, Haocheng, et autres
Publié: (2026)
par: Xi, Haocheng, et autres
Publié: (2026)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
par: Yu, Zhongzhi, et autres
Publié: (2024)
par: Yu, Zhongzhi, et autres
Publié: (2024)
Perturbed examples reveal invariances shared by language models
par: Rawal, Ruchit, et autres
Publié: (2023)
par: Rawal, Ruchit, et autres
Publié: (2023)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
par: Wong, Liang Ze
Publié: (2025)
par: Wong, Liang Ze
Publié: (2025)
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
par: Wang, Dingzirui, et autres
Publié: (2025)
par: Wang, Dingzirui, et autres
Publié: (2025)
Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
par: Roy, Debjyoti Saha, et autres
Publié: (2024)
par: Roy, Debjyoti Saha, et autres
Publié: (2024)
ProxyAttn: Guided Sparse Attention via Representative Heads
par: Wang, Yixuan, et autres
Publié: (2025)
par: Wang, Yixuan, et autres
Publié: (2025)
Documents similaires
-
Controlling changes to attention logits
par: Anson, Ben, et autres
Publié: (2025) -
Learning to Skip the Middle Layers of Transformers
par: Lawson, Tim, et autres
Publié: (2025) -
Batch size invariant Adam
par: Wang, Xi, et autres
Publié: (2024) -
Function-Space Learning Rates
par: Milsom, Edward, et autres
Publié: (2025) -
Flexible Infinite-Width Graph Convolutional Neural Networks
par: Anson, Ben, et autres
Publié: (2024)