The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
Fuente:
arXiv
Salvato in:
| Autore principale: | Balogh, Peter |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Do Neurons Dream of Primitive Operators? Wake-Sleep Compression Rediscovers Schank's Event Semantics
di: Balogh, Peter
Pubblicazione: (2026)
di: Balogh, Peter
Pubblicazione: (2026)
RecurFormer: Not All Transformer Heads Need Self-Attention
di: Yan, Ruiqing, et al.
Pubblicazione: (2024)
di: Yan, Ruiqing, et al.
Pubblicazione: (2024)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
di: Chen, Yilong, et al.
Pubblicazione: (2024)
di: Chen, Yilong, et al.
Pubblicazione: (2024)
Which Attention Heads Matter for In-Context Learning?
di: Yin, Kayo, et al.
Pubblicazione: (2025)
di: Yin, Kayo, et al.
Pubblicazione: (2025)
FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
Selective Attention Improves Transformer
di: Leviathan, Yaniv, et al.
Pubblicazione: (2024)
di: Leviathan, Yaniv, et al.
Pubblicazione: (2024)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
di: Mąka, Paweł, et al.
Pubblicazione: (2024)
di: Mąka, Paweł, et al.
Pubblicazione: (2024)
What Matters in Transformers? Not All Attention is Needed
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
Forgetting Transformer: Softmax Attention with a Forget Gate
di: Lin, Zhixuan, et al.
Pubblicazione: (2025)
di: Lin, Zhixuan, et al.
Pubblicazione: (2025)
On the Role of Attention Heads in Large Language Model Safety
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
di: Collins, Liam, et al.
Pubblicazione: (2024)
di: Collins, Liam, et al.
Pubblicazione: (2024)
Pooling Attention: Evaluating Pretrained Transformer Embeddings for Deception Classification
di: Mamtani, Sumit, et al.
Pubblicazione: (2025)
di: Mamtani, Sumit, et al.
Pubblicazione: (2025)
Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
di: Koh, Minsu, et al.
Pubblicazione: (2025)
di: Koh, Minsu, et al.
Pubblicazione: (2025)
Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
di: Choi, Minsik, et al.
Pubblicazione: (2025)
di: Choi, Minsik, et al.
Pubblicazione: (2025)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
di: Adhikari, Rabin
Pubblicazione: (2025)
di: Adhikari, Rabin
Pubblicazione: (2025)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
di: Su, Jingtong, et al.
Pubblicazione: (2025)
di: Su, Jingtong, et al.
Pubblicazione: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
di: Neo, Clement, et al.
Pubblicazione: (2024)
di: Neo, Clement, et al.
Pubblicazione: (2024)
Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget
di: Balogh, Peter
Pubblicazione: (2026)
di: Balogh, Peter
Pubblicazione: (2026)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
di: Leemann, Tobias, et al.
Pubblicazione: (2024)
di: Leemann, Tobias, et al.
Pubblicazione: (2024)
Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention
di: Guo, Zhenyu, et al.
Pubblicazione: (2025)
di: Guo, Zhenyu, et al.
Pubblicazione: (2025)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
di: Tian, Yuandong, et al.
Pubblicazione: (2023)
di: Tian, Yuandong, et al.
Pubblicazione: (2023)
Multi-Head Attention Is a Multi-Player Game
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2026)
di: Chakrabarti, Kushal, et al.
Pubblicazione: (2026)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
di: Zhussip, Magauiya, et al.
Pubblicazione: (2025)
di: Zhussip, Magauiya, et al.
Pubblicazione: (2025)
GLU Attention Improve Transformer
di: Wang, Zehao
Pubblicazione: (2025)
di: Wang, Zehao
Pubblicazione: (2025)
Multi-Head Mixture-of-Experts
di: Wu, Xun, et al.
Pubblicazione: (2024)
di: Wu, Xun, et al.
Pubblicazione: (2024)
MAGE: Multi-Head Attention Guided Embeddings for Low Resource Sentiment Classification
di: Vashisht, Varun, et al.
Pubblicazione: (2025)
di: Vashisht, Varun, et al.
Pubblicazione: (2025)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
di: Wang, Hanrui, et al.
Pubblicazione: (2020)
di: Wang, Hanrui, et al.
Pubblicazione: (2020)
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
di: Mistry, Deven Mahesh, et al.
Pubblicazione: (2025)
di: Mistry, Deven Mahesh, et al.
Pubblicazione: (2025)
Cultural Binding Heads in Language Models
di: Floro, Avrile, et al.
Pubblicazione: (2026)
di: Floro, Avrile, et al.
Pubblicazione: (2026)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
di: Zhu, Shiyi, et al.
Pubblicazione: (2023)
di: Zhu, Shiyi, et al.
Pubblicazione: (2023)
Filtered Direct Preference Optimization
di: Morimura, Tetsuro, et al.
Pubblicazione: (2024)
di: Morimura, Tetsuro, et al.
Pubblicazione: (2024)
EcoTransformer: Attention without Multiplication
di: Gao, Xin, et al.
Pubblicazione: (2025)
di: Gao, Xin, et al.
Pubblicazione: (2025)
Efficient Systematic Reviews: Literature Filtering with Transformers & Transfer Learning
di: Hawkins, John, et al.
Pubblicazione: (2024)
di: Hawkins, John, et al.
Pubblicazione: (2024)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
di: Chen, Siyu, et al.
Pubblicazione: (2024)
di: Chen, Siyu, et al.
Pubblicazione: (2024)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
di: Constantinou, Christos, et al.
Pubblicazione: (2024)
di: Constantinou, Christos, et al.
Pubblicazione: (2024)
Iteration Head: A Mechanistic Study of Chain-of-Thought
di: Cabannes, Vivien, et al.
Pubblicazione: (2024)
di: Cabannes, Vivien, et al.
Pubblicazione: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
di: Huang, Yuxiang, et al.
Pubblicazione: (2026)
di: Huang, Yuxiang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Do Neurons Dream of Primitive Operators? Wake-Sleep Compression Rediscovers Schank's Event Semantics
di: Balogh, Peter
Pubblicazione: (2026) -
RecurFormer: Not All Transformer Heads Need Self-Attention
di: Yan, Ruiqing, et al.
Pubblicazione: (2024) -
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
di: Chen, Yilong, et al.
Pubblicazione: (2024) -
Which Attention Heads Matter for In-Context Learning?
di: Yin, Kayo, et al.
Pubblicazione: (2025) -
FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning
di: Luo, Haozheng, et al.
Pubblicazione: (2026)