RecurFormer: Not All Transformer Heads Need Self-Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Ruiqing, Zheng, Linghan, Du, Xingbo, Zou, Han, Guo, Yufeng, Yang, Jianfei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unveiling and Controlling Anomalous Attention Distribution in Transformers
by: Yan, Ruiqing, et al.
Published: (2024)
by: Yan, Ruiqing, et al.
Published: (2024)
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Tensor Product Attention Is All You Need
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026)
by: Zade, Saleh Zare, et al.
Published: (2026)
Attention Is All You Need for KV Cache in Diffusion LLMs
by: Nguyen-Tri, Quan, et al.
Published: (2025)
by: Nguyen-Tri, Quan, et al.
Published: (2025)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
by: Balogh, Peter
Published: (2026)
by: Balogh, Peter
Published: (2026)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
by: Fu, Zichuan, et al.
Published: (2026)
by: Fu, Zichuan, et al.
Published: (2026)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
Forget Attention: Importance-Aware Attention Is All You Need
by: Shin, Soohyeong, et al.
Published: (2026)
by: Shin, Soohyeong, et al.
Published: (2026)
Synthetic Data RL: Task Definition Is All You Need
by: Guo, Yiduo, et al.
Published: (2025)
by: Guo, Yiduo, et al.
Published: (2025)
All You Need is One: Capsule Prompt Tuning with a Single Vector
by: Liu, Yiyang, et al.
Published: (2025)
by: Liu, Yiyang, et al.
Published: (2025)
TransMLA: Multi-Head Latent Attention Is All You Need
by: Meng, Fanxu, et al.
Published: (2025)
by: Meng, Fanxu, et al.
Published: (2025)
One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels
by: Seshadri, Amrit Diggavi
Published: (2025)
by: Seshadri, Amrit Diggavi
Published: (2025)
More Agents Is All You Need
by: Li, Junyou, et al.
Published: (2024)
by: Li, Junyou, et al.
Published: (2024)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
by: Ildiz, M. Emrullah, et al.
Published: (2024)
by: Ildiz, M. Emrullah, et al.
Published: (2024)
ToolRL: Reward is All Tool Learning Needs
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025)
by: Hiranandani, Gaurush, et al.
Published: (2025)
IoT-LLM: a framework for enhancing Large Language Model reasoning from real-world sensor data
by: An, Tuo, et al.
Published: (2024)
by: An, Tuo, et al.
Published: (2024)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention
by: Guo, Zhenyu, et al.
Published: (2025)
by: Guo, Zhenyu, et al.
Published: (2025)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
by: Tian, Yuandong, et al.
Published: (2023)
by: Tian, Yuandong, et al.
Published: (2023)
Think When You Need: Self-Adaptive Chain-of-Thought Learning
by: Yang, Junjie, et al.
Published: (2025)
by: Yang, Junjie, et al.
Published: (2025)
Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing
by: Cong, Peizhuang, et al.
Published: (2024)
by: Cong, Peizhuang, et al.
Published: (2024)
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
by: Gan, Chunjing, et al.
Published: (2024)
by: Gan, Chunjing, et al.
Published: (2024)
On the Role of Attention Heads in Large Language Model Safety
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Language is All a Graph Needs
by: Ye, Ruosong, et al.
Published: (2023)
by: Ye, Ruosong, et al.
Published: (2023)
Selective Attention Improves Transformer
by: Leviathan, Yaniv, et al.
Published: (2024)
by: Leviathan, Yaniv, et al.
Published: (2024)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
NoiseFormer -- Noise Diffused Symmetric Attention Transformer
by: Kumar, Phani, et al.
Published: (2026)
by: Kumar, Phani, et al.
Published: (2026)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
by: Chen, Lingjiao, et al.
Published: (2024)
by: Chen, Lingjiao, et al.
Published: (2024)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
by: Steinmetz, Cody, et al.
Published: (2025)
by: Steinmetz, Cody, et al.
Published: (2025)
Guidance is All You Need: Temperature-Guided Reasoning in Large Language Models
by: Gomaa, Eyad, et al.
Published: (2024)
by: Gomaa, Eyad, et al.
Published: (2024)
More Compute Is What You Need
by: Guo, Zhen
Published: (2024)
by: Guo, Zhen
Published: (2024)
From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary
by: Zheng, Qirui, et al.
Published: (2025)
by: Zheng, Qirui, et al.
Published: (2025)
All Language Models Large and Small
by: Chen, Zhixun, et al.
Published: (2024)
by: Chen, Zhixun, et al.
Published: (2024)
Attention is All You Need Until You Need Retention
by: Yaslioglu, M. Murat
Published: (2025)
by: Yaslioglu, M. Murat
Published: (2025)
Demystifying the Slash Pattern in Attention: The Role of RoPE
by: Cheng, Yuan, et al.
Published: (2026)
by: Cheng, Yuan, et al.
Published: (2026)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025)
by: Lin, Feng, et al.
Published: (2025)
Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
by: Karbevski, Marko, et al.
Published: (2025)
by: Karbevski, Marko, et al.
Published: (2025)
Similar Items
-
Unveiling and Controlling Anomalous Attention Distribution in Transformers
by: Yan, Ruiqing, et al.
Published: (2024) -
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024) -
Tensor Product Attention Is All You Need
by: Zhang, Yifan, et al.
Published: (2025) -
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026) -
Attention Is All You Need for KV Cache in Diffusion LLMs
by: Nguyen-Tri, Quan, et al.
Published: (2025)