Attention Needs to Focus: A Unified Perspective on Attention Allocation
Fuente:
arXiv
Salvato in:
| Autori principali: | Fu, Zichuan, Song, Wentao, Li, Guojing, Wang, Yejing, Wu, Xian, Deng, Yimin, Yan, Hanyu, Zheng, Yefeng, Zhao, Xiangyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Sliding Window Attention Training for Efficient Large Language Models
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
Model Merging for Knowledge Editing
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models
di: Deng, Yimin, et al.
Pubblicazione: (2026)
di: Deng, Yimin, et al.
Pubblicazione: (2026)
Tandem: Riding Together with Large and Small Language Models for Efficient Reasoning
di: Fu, Zichuan, et al.
Pubblicazione: (2026)
di: Fu, Zichuan, et al.
Pubblicazione: (2026)
MultiDx: A Multi-Source Knowledge Integration Framework towards Diagnostic Reasoning
di: Deng, Yimin, et al.
Pubblicazione: (2026)
di: Deng, Yimin, et al.
Pubblicazione: (2026)
Training-free LLM Merging for Multi-task Learning
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
How Large Language Models Need Symbolism
di: Deng, Xiaotie, et al.
Pubblicazione: (2025)
di: Deng, Xiaotie, et al.
Pubblicazione: (2025)
Attention Smoothing Is All You Need For Unlearning
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
When MOE Meets LLMs: Parameter Efficient Fine-tuning for Multi-task Medical Applications
di: Liu, Qidong, et al.
Pubblicazione: (2023)
di: Liu, Qidong, et al.
Pubblicazione: (2023)
MedKP: Medical Dialogue with Knowledge Enhancement and Clinical Pathway Encoding
di: Wu, Jiageng, et al.
Pubblicazione: (2024)
di: Wu, Jiageng, et al.
Pubblicazione: (2024)
A Multi-Expert Structural-Semantic Hybrid Framework for Unveiling Historical Patterns in Temporal Knowledge Graphs
di: Deng, Yimin, et al.
Pubblicazione: (2025)
di: Deng, Yimin, et al.
Pubblicazione: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
RecurFormer: Not All Transformer Heads Need Self-Attention
di: Yan, Ruiqing, et al.
Pubblicazione: (2024)
di: Yan, Ruiqing, et al.
Pubblicazione: (2024)
Contextual Attention Modulation: Towards Efficient Multi-Task Adaptation in Large Language Models
di: Pan, Dayan, et al.
Pubblicazione: (2025)
di: Pan, Dayan, et al.
Pubblicazione: (2025)
Biomedical Entity Linking as Multiple Choice Question Answering
di: Lin, Zhenxi, et al.
Pubblicazione: (2024)
di: Lin, Zhenxi, et al.
Pubblicazione: (2024)
Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data
di: Guan, Zhong, et al.
Pubblicazione: (2025)
di: Guan, Zhong, et al.
Pubblicazione: (2025)
Neural Attention Search
di: Deng, Difan, et al.
Pubblicazione: (2025)
di: Deng, Difan, et al.
Pubblicazione: (2025)
Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models
di: Liu, Xin, et al.
Pubblicazione: (2025)
di: Liu, Xin, et al.
Pubblicazione: (2025)
SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
di: Li, Mengjie, et al.
Pubblicazione: (2025)
di: Li, Mengjie, et al.
Pubblicazione: (2025)
ReAttention: Training-Free Infinite Context with Finite Attention Scope
di: Liu, Xiaoran, et al.
Pubblicazione: (2024)
di: Liu, Xiaoran, et al.
Pubblicazione: (2024)
Pay Attention to What You Need
di: Gao, Yifei, et al.
Pubblicazione: (2023)
di: Gao, Yifei, et al.
Pubblicazione: (2023)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
di: Yun, Jungmin, et al.
Pubblicazione: (2024)
di: Yun, Jungmin, et al.
Pubblicazione: (2024)
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
What Matters in Transformers? Not All Attention is Needed
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
GTA: Grouped-head latenT Attention
di: Sun, Luoyang, et al.
Pubblicazione: (2025)
di: Sun, Luoyang, et al.
Pubblicazione: (2025)
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
di: Li, Shuaiyi, et al.
Pubblicazione: (2026)
di: Li, Shuaiyi, et al.
Pubblicazione: (2026)
Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency
di: Zhao, Hanyu, et al.
Pubblicazione: (2024)
di: Zhao, Hanyu, et al.
Pubblicazione: (2024)
Editing Factual Knowledge and Explanatory Ability of Medical Large Language Models
di: Xu, Derong, et al.
Pubblicazione: (2024)
di: Xu, Derong, et al.
Pubblicazione: (2024)
Scaling Reasoning without Attention
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)
di: Zhao, Xueliang, et al.
Pubblicazione: (2025)
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
di: Yang, Ning, et al.
Pubblicazione: (2026)
di: Yang, Ning, et al.
Pubblicazione: (2026)
Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens
di: Yong, Xixian, et al.
Pubblicazione: (2025)
di: Yong, Xixian, et al.
Pubblicazione: (2025)
Multi-Perspective Attention Mechanism for Bias-Aware Sequential Recommendation
di: Fu, Mingjian, et al.
Pubblicazione: (2025)
di: Fu, Mingjian, et al.
Pubblicazione: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
di: Nguyen-Tri, Quan, et al.
Pubblicazione: (2025)
di: Nguyen-Tri, Quan, et al.
Pubblicazione: (2025)
HSR-Enhanced Sparse Attention Acceleration
di: Chen, Bo, et al.
Pubblicazione: (2024)
di: Chen, Bo, et al.
Pubblicazione: (2024)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
di: Yang, Lijie, et al.
Pubblicazione: (2025)
di: Yang, Lijie, et al.
Pubblicazione: (2025)
HiCI: Hierarchical Construction-Integration for Long-Context Attention
di: Zeng, Xiangyu, et al.
Pubblicazione: (2026)
di: Zeng, Xiangyu, et al.
Pubblicazione: (2026)
Efficient Streaming Language Models with Attention Sinks
di: Xiao, Guangxuan, et al.
Pubblicazione: (2023)
di: Xiao, Guangxuan, et al.
Pubblicazione: (2023)
Multi-granularity Interactive Attention Framework for Residual Hierarchical Pronunciation Assessment
di: Han, Hong, et al.
Pubblicazione: (2026)
di: Han, Hong, et al.
Pubblicazione: (2026)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
di: Xu, Yixing, et al.
Pubblicazione: (2025)
di: Xu, Yixing, et al.
Pubblicazione: (2025)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
di: Bae, Jeongin, et al.
Pubblicazione: (2026)
di: Bae, Jeongin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Sliding Window Attention Training for Efficient Large Language Models
di: Fu, Zichuan, et al.
Pubblicazione: (2025) -
Model Merging for Knowledge Editing
di: Fu, Zichuan, et al.
Pubblicazione: (2025) -
AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models
di: Deng, Yimin, et al.
Pubblicazione: (2026) -
Tandem: Riding Together with Large and Small Language Models for Efficient Reasoning
di: Fu, Zichuan, et al.
Pubblicazione: (2026) -
MultiDx: A Multi-Source Knowledge Integration Framework towards Diagnostic Reasoning
di: Deng, Yimin, et al.
Pubblicazione: (2026)