Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Difan, Winje, Andreas Bentzen, Fehring, Lukas, Lindauer, Marius |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neural Attention Search
von: Deng, Difan, et al.
Veröffentlicht: (2025)
von: Deng, Difan, et al.
Veröffentlicht: (2025)
Optimizing Time Series Forecasting Architectures: A Hierarchical Neural Architecture Search Approach
von: Deng, Difan, et al.
Veröffentlicht: (2024)
von: Deng, Difan, et al.
Veröffentlicht: (2024)
Growing with Experience: Growing Neural Networks in Deep Reinforcement Learning
von: Fehring, Lukas, et al.
Veröffentlicht: (2025)
von: Fehring, Lukas, et al.
Veröffentlicht: (2025)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
von: He, Mutian, et al.
Veröffentlicht: (2025)
von: He, Mutian, et al.
Veröffentlicht: (2025)
Unifying Linear-Time Attention via Latent Probabilistic Modelling
von: Dolga, Rares, et al.
Veröffentlicht: (2024)
von: Dolga, Rares, et al.
Veröffentlicht: (2024)
Why Softmax Attention Outperforms Linear Attention
von: Deng, Yichuan, et al.
Veröffentlicht: (2023)
von: Deng, Yichuan, et al.
Veröffentlicht: (2023)
AutoML in the Age of Large Language Models: Current Challenges, Future Opportunities and Risks
von: Tornede, Alexander, et al.
Veröffentlicht: (2023)
von: Tornede, Alexander, et al.
Veröffentlicht: (2023)
Nectar: Neural Estimation of Cached-Token Attention via Regression
von: Monteiro, João, et al.
Veröffentlicht: (2026)
von: Monteiro, João, et al.
Veröffentlicht: (2026)
SEA: Sparse Linear Attention with Estimated Attention Mask
von: Lee, Heejun, et al.
Veröffentlicht: (2023)
von: Lee, Heejun, et al.
Veröffentlicht: (2023)
Dynamic Priors in Bayesian Optimization for Hyperparameter Optimization
von: Fehring, Lukas, et al.
Veröffentlicht: (2025)
von: Fehring, Lukas, et al.
Veröffentlicht: (2025)
Token Distillation: Attention-aware Input Embeddings For New Tokens
von: Dobler, Konstantin, et al.
Veröffentlicht: (2025)
von: Dobler, Konstantin, et al.
Veröffentlicht: (2025)
Linear Attention Sequence Parallelism
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
Softmax Attention with Constant Cost per Token
von: Heinsen, Franz A.
Veröffentlicht: (2024)
von: Heinsen, Franz A.
Veröffentlicht: (2024)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2025)
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2025)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
von: Jin, Zehao, et al.
Veröffentlicht: (2026)
von: Jin, Zehao, et al.
Veröffentlicht: (2026)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
STS: Efficient Sparse Attention with Speculative Token Sparsity
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
Attention with Trained Embeddings Provably Selects Important Tokens
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Scaling Linear Attention with Sparse State Expansion
von: Pan, Yuqi, et al.
Veröffentlicht: (2025)
von: Pan, Yuqi, et al.
Veröffentlicht: (2025)
Kimi Linear: An Expressive, Efficient Attention Architecture
von: Kimi Team, et al.
Veröffentlicht: (2025)
von: Kimi Team, et al.
Veröffentlicht: (2025)
Towards Token-Level Text Anomaly Detection
von: Cao, Yang, et al.
Veröffentlicht: (2026)
von: Cao, Yang, et al.
Veröffentlicht: (2026)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
von: You, Haoran, et al.
Veröffentlicht: (2024)
von: You, Haoran, et al.
Veröffentlicht: (2024)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2026)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2026)
S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models
von: Young, Jack
Veröffentlicht: (2026)
von: Young, Jack
Veröffentlicht: (2026)
Higher-order Linear Attention
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Native Hybrid Attention for Efficient Sequence Modeling
von: Du, Jusen, et al.
Veröffentlicht: (2025)
von: Du, Jusen, et al.
Veröffentlicht: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
von: Mihaila, George
Veröffentlicht: (2026)
von: Mihaila, George
Veröffentlicht: (2026)
Gated Linear Attention Transformers with Hardware-Efficient Training
von: Yang, Songlin, et al.
Veröffentlicht: (2023)
von: Yang, Songlin, et al.
Veröffentlicht: (2023)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
von: Zhang, Michael, et al.
Veröffentlicht: (2024)
von: Zhang, Michael, et al.
Veröffentlicht: (2024)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
von: De, Soham, et al.
Veröffentlicht: (2024)
von: De, Soham, et al.
Veröffentlicht: (2024)
AdaSplash: Adaptive Sparse Flash Attention
von: Gonçalves, Nuno, et al.
Veröffentlicht: (2025)
von: Gonçalves, Nuno, et al.
Veröffentlicht: (2025)
Parallax: Parameterized Local Linear Attention for Language Modeling
von: Zuo, Yifei, et al.
Veröffentlicht: (2026)
von: Zuo, Yifei, et al.
Veröffentlicht: (2026)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
von: Bozic, Vukasin, et al.
Veröffentlicht: (2023)
von: Bozic, Vukasin, et al.
Veröffentlicht: (2023)
DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU
von: Chen, Weizhe, et al.
Veröffentlicht: (2026)
von: Chen, Weizhe, et al.
Veröffentlicht: (2026)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
von: Nagaraj, Manish, et al.
Veröffentlicht: (2025)
von: Nagaraj, Manish, et al.
Veröffentlicht: (2025)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Neural Attention Search
von: Deng, Difan, et al.
Veröffentlicht: (2025) -
Optimizing Time Series Forecasting Architectures: A Hierarchical Neural Architecture Search Approach
von: Deng, Difan, et al.
Veröffentlicht: (2024) -
Growing with Experience: Growing Neural Networks in Deep Reinforcement Learning
von: Fehring, Lukas, et al.
Veröffentlicht: (2025) -
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
von: He, Mutian, et al.
Veröffentlicht: (2025) -
Unifying Linear-Time Attention via Latent Probabilistic Modelling
von: Dolga, Rares, et al.
Veröffentlicht: (2024)