Learning to Focus: Focal Attention for Selective and Scalable Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Ram, Dhananjay, Xia, Wei, Soatto, Stefano |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models
by: Liu, Xiaoze, et al.
Published: (2026)
by: Liu, Xiaoze, et al.
Published: (2026)
Cycles of Thought: Measuring LLM Confidence through Stable Explanations
by: Becker, Evan, et al.
Published: (2024)
by: Becker, Evan, et al.
Published: (2024)
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
by: Choudhary, Sakshi, et al.
Published: (2026)
by: Choudhary, Sakshi, et al.
Published: (2026)
Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression
by: Peng, Liangzu, et al.
Published: (2025)
by: Peng, Liangzu, et al.
Published: (2025)
Expansion Span: Combining Fading Memory and Retrieval in Hybrid State Space Models
by: Nunez, Elvis, et al.
Published: (2024)
by: Nunez, Elvis, et al.
Published: (2024)
Selective Attention Improves Transformer
by: Leviathan, Yaniv, et al.
Published: (2024)
by: Leviathan, Yaniv, et al.
Published: (2024)
Selective Attention: Enhancing Transformer through Principled Context Control
by: Zhang, Xuechen, et al.
Published: (2024)
by: Zhang, Xuechen, et al.
Published: (2024)
Adaptive Two Sided Laplace Transforms: A Learnable, Interpretable, and Scalable Replacement for Self-Attention
by: Kiruluta, Andrew
Published: (2025)
by: Kiruluta, Andrew
Published: (2025)
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
Meanings and Feelings of Large Language Models: Observability of Latent States in Generative AI
by: Liu, Tian Yu, et al.
Published: (2024)
by: Liu, Tian Yu, et al.
Published: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Fast Multipole Attention: A Scalable Multilevel Attention Mechanism for Text and Images
by: Kang, Yanming, et al.
Published: (2023)
by: Kang, Yanming, et al.
Published: (2023)
Sequence-level Large Language Model Training with Contrastive Preference Optimization
by: Feng, Zhili, et al.
Published: (2025)
by: Feng, Zhili, et al.
Published: (2025)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
by: Fu, Zichuan, et al.
Published: (2026)
by: Fu, Zichuan, et al.
Published: (2026)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
by: Mihaila, George
Published: (2026)
by: Mihaila, George
Published: (2026)
LASER: Attention with Exponential Transformation
by: Duvvuri, Sai Surya, et al.
Published: (2024)
by: Duvvuri, Sai Surya, et al.
Published: (2024)
Says Who? Effective Zero-Shot Annotation of Focalization
by: Hicke, Rebecca M. M., et al.
Published: (2024)
by: Hicke, Rebecca M. M., et al.
Published: (2024)
HLAT: High-quality Large Language Model Pre-trained on AWS Trainium
by: Fan, Haozheng, et al.
Published: (2024)
by: Fan, Haozheng, et al.
Published: (2024)
Language Models Can Predict Their Own Behavior
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
A Little Human Data Goes A Long Way
by: Ashok, Dhananjay, et al.
Published: (2024)
by: Ashok, Dhananjay, et al.
Published: (2024)
PICASO: Permutation-Invariant Context Composition with State Space Models
by: Liu, Tian Yu, et al.
Published: (2025)
by: Liu, Tian Yu, et al.
Published: (2025)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025)
by: Wang, Junxuan, et al.
Published: (2025)
Generalized Probabilistic Attention Mechanism in Transformers
by: Heo, DongNyeong, et al.
Published: (2024)
by: Heo, DongNyeong, et al.
Published: (2024)
SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language Models
by: He, Jinghan, et al.
Published: (2024)
by: He, Jinghan, et al.
Published: (2024)
Fewer Truncations Improve Language Modeling
by: Ding, Hantian, et al.
Published: (2024)
by: Ding, Hantian, et al.
Published: (2024)
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
by: Aggarwal, Shubham, et al.
Published: (2026)
by: Aggarwal, Shubham, et al.
Published: (2026)
Teaching Large Language Models Number-Focused Headline Generation With Key Element Rationales
by: Qian, Zhen, et al.
Published: (2025)
by: Qian, Zhen, et al.
Published: (2025)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
Differential Transformer
by: Ye, Tianzhu, et al.
Published: (2024)
by: Ye, Tianzhu, et al.
Published: (2024)
Attention with Trained Embeddings Provably Selects Important Tokens
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
Do Sentence Transformers Learn Quasi-Geospatial Concepts from General Text?
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
Extracting Rule-based Descriptions of Attention Features in Transformers
by: Friedman, Dan, et al.
Published: (2025)
by: Friedman, Dan, et al.
Published: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
by: Xiao, Da, et al.
Published: (2024)
by: Xiao, Da, et al.
Published: (2024)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Transformers with Selective Access to Early Representations
by: Gunasekaran, Skye, et al.
Published: (2026)
by: Gunasekaran, Skye, et al.
Published: (2026)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
by: Xiao, Kaicheng, et al.
Published: (2026)
by: Xiao, Kaicheng, et al.
Published: (2026)
Learning to Attribute with Attention
by: Cohen-Wang, Benjamin, et al.
Published: (2025)
by: Cohen-Wang, Benjamin, et al.
Published: (2025)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024)
by: Collins, Liam, et al.
Published: (2024)
Similar Items
-
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models
by: Liu, Xiaoze, et al.
Published: (2026) -
Cycles of Thought: Measuring LLM Confidence through Stable Explanations
by: Becker, Evan, et al.
Published: (2024) -
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
by: Choudhary, Sakshi, et al.
Published: (2026) -
Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression
by: Peng, Liangzu, et al.
Published: (2025) -
Expansion Span: Combining Fading Memory and Retrieval in Hybrid State Space Models
by: Nunez, Elvis, et al.
Published: (2024)