Why Attend to Everything? Focus is the Key
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Hengshuai, Chen, Xing, Murtadha, Ahmed, Li, Jin, Yadkori, Yasin Abbasi, Shao, Shuai, Liu, Changling, Wang, Guan, Yuan, Mingli, Chen, William, Song, Sen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thin Keys, Full Values: Reducing KV Cache via Low-Dimensional Attention Selection
by: Yao, Hengshuai, et al.
Published: (2026)
by: Yao, Hengshuai, et al.
Published: (2026)
GAIN: Multiplicative Modulation for Domain Adaptation
by: Yao, Hengshuai, et al.
Published: (2026)
by: Yao, Hengshuai, et al.
Published: (2026)
HRM-Text: Efficient Pretraining Beyond Scaling
by: Wang, Guan, et al.
Published: (2026)
by: Wang, Guan, et al.
Published: (2026)
Hierarchical Reasoning Model
by: Wang, Guan, et al.
Published: (2025)
by: Wang, Guan, et al.
Published: (2025)
To Believe or Not to Believe Your LLM
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
Pointwise confidence estimation in the non-linear $\ell^2$-regularized least squares
by: Kuzborskij, Ilja, et al.
Published: (2025)
by: Kuzborskij, Ilja, et al.
Published: (2025)
Low-rank bias, weight decay, and model merging in neural networks
by: Kuzborskij, Ilja, et al.
Published: (2025)
by: Kuzborskij, Ilja, et al.
Published: (2025)
Supervised Gradual Machine Learning for Aspect Category Detection
by: Ahmed, Murtadha, et al.
Published: (2024)
by: Ahmed, Murtadha, et al.
Published: (2024)
MateICL: Mitigating Attention Dispersion in Large-Scale In-Context Learning
by: Ahmed, Murtadha, et al.
Published: (2025)
by: Ahmed, Murtadha, et al.
Published: (2025)
Mitigating LLM Hallucinations via Conformal Abstention
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
AlcLaM: Arabic Dialectal Language Model
by: Ahmed, Murtadha, et al.
Published: (2024)
by: Ahmed, Murtadha, et al.
Published: (2024)
Naive Bayes-based Context Extension for Large Language Models
by: Su, Jianlin, et al.
Published: (2024)
by: Su, Jianlin, et al.
Published: (2024)
Best of both worlds: Stochastic & adversarial best-arm identification
by: Abbasi-Yadkori, Yasin, et al.
Published: (2026)
by: Abbasi-Yadkori, Yasin, et al.
Published: (2026)
Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
by: Chen, Shiqi, et al.
Published: (2025)
by: Chen, Shiqi, et al.
Published: (2025)
HTAM: Hierarchical Transition-Attended Memory for Operator Optimization
by: Zhang, Yining, et al.
Published: (2026)
by: Zhang, Yining, et al.
Published: (2026)
PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
by: Chen, Sihan, et al.
Published: (2025)
by: Chen, Sihan, et al.
Published: (2025)
BERT-ASC: Auxiliary-Sentence Construction for Implicit Aspect Learning in Sentiment Analysis
by: Ahmed, Murtadha, et al.
Published: (2022)
by: Ahmed, Murtadha, et al.
Published: (2022)
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
by: Choi, Younwoo, et al.
Published: (2025)
by: Choi, Younwoo, et al.
Published: (2025)
Learning When Not to Attend Globally
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
Deep CLAS: Deep Contextual Listen, Attend and Spell
by: Wang, Mengzhi, et al.
Published: (2024)
by: Wang, Mengzhi, et al.
Published: (2024)
PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
by: Tang, Yixuan, et al.
Published: (2025)
by: Tang, Yixuan, et al.
Published: (2025)
The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
by: Qian, Chen, et al.
Published: (2026)
by: Qian, Chen, et al.
Published: (2026)
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
by: Hu, Lingxiang, et al.
Published: (2025)
by: Hu, Lingxiang, et al.
Published: (2025)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
by: Lyu, Zhiheng, et al.
Published: (2025)
by: Lyu, Zhiheng, et al.
Published: (2025)
Don't Lose Focus: Activation Steering via Key-Orthogonal Projections
by: Luo, Haoyan, et al.
Published: (2026)
by: Luo, Haoyan, et al.
Published: (2026)
Attamba: Attending To Multi-Token States
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model Alignment
by: Chen, Zhipeng, et al.
Published: (2024)
by: Chen, Zhipeng, et al.
Published: (2024)
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
by: Rajabzadeh, Hossein, et al.
Published: (2024)
by: Rajabzadeh, Hossein, et al.
Published: (2024)
Causal Attention with Lookahead Keys
by: Song, Zhuoqing, et al.
Published: (2025)
by: Song, Zhuoqing, et al.
Published: (2025)
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
by: Hu, Ruina, et al.
Published: (2026)
by: Hu, Ruina, et al.
Published: (2026)
Timing is Everything: Temporal Scaffolding of Semantic Surprise in Humor
by: Ma, Yuxi, et al.
Published: (2026)
by: Ma, Yuxi, et al.
Published: (2026)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Benchmarking Sociolinguistic Diversity in Swahili NLP: A Taxonomy-Guided Approach
by: Oketch, Kezia, et al.
Published: (2025)
by: Oketch, Kezia, et al.
Published: (2025)
Towards Efficient LLM-aware Heterogeneous Graph Learning
by: Li, Wenda, et al.
Published: (2025)
by: Li, Wenda, et al.
Published: (2025)
Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information
by: Chen, Yao, et al.
Published: (2026)
by: Chen, Yao, et al.
Published: (2026)
Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers
by: Ben-Artzy, Amit, et al.
Published: (2024)
by: Ben-Artzy, Amit, et al.
Published: (2024)
Bridging the LLM Accessibility Divide? Performance, Fairness, and Cost of Closed versus Open LLMs for Automated Essay Scoring
by: Oketch, Kezia, et al.
Published: (2025)
by: Oketch, Kezia, et al.
Published: (2025)
Decomposing and Measuring Evaluation Awareness
by: Li, Changling, et al.
Published: (2026)
by: Li, Changling, et al.
Published: (2026)
Why Attention Patterns Exist: A Unifying Temporal Perspective Analysis
by: Yang, Qingyue, et al.
Published: (2026)
by: Yang, Qingyue, et al.
Published: (2026)
Similar Items
-
Thin Keys, Full Values: Reducing KV Cache via Low-Dimensional Attention Selection
by: Yao, Hengshuai, et al.
Published: (2026) -
GAIN: Multiplicative Modulation for Domain Adaptation
by: Yao, Hengshuai, et al.
Published: (2026) -
HRM-Text: Efficient Pretraining Beyond Scaling
by: Wang, Guan, et al.
Published: (2026) -
Hierarchical Reasoning Model
by: Wang, Guan, et al.
Published: (2025) -
To Believe or Not to Believe Your LLM
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)