Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Jin, Zehao, Sui, Yanan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Kimi Linear: An Expressive, Efficient Attention Architecture
di: Kimi Team, et al.
Pubblicazione: (2025)
di: Kimi Team, et al.
Pubblicazione: (2025)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
di: Zhang, Michael, et al.
Pubblicazione: (2024)
di: Zhang, Michael, et al.
Pubblicazione: (2024)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
di: Xiao, Kaicheng, et al.
Pubblicazione: (2026)
di: Xiao, Kaicheng, et al.
Pubblicazione: (2026)
GLU Attention Improve Transformer
di: Wang, Zehao
Pubblicazione: (2025)
di: Wang, Zehao
Pubblicazione: (2025)
Routing Absorption in Sparse Attention: Why Random Gates Are Hard to Beat
di: Aquino-Michaels, Keston
Pubblicazione: (2026)
di: Aquino-Michaels, Keston
Pubblicazione: (2026)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
di: Dentamaro, Vincenzo
Pubblicazione: (2025)
di: Dentamaro, Vincenzo
Pubblicazione: (2025)
Whole-Brain Connectomic Graph Model Enables Whole-Body Locomotion Control in Fruit Fly
di: Jin, Zehao, et al.
Pubblicazione: (2026)
di: Jin, Zehao, et al.
Pubblicazione: (2026)
More Expressive Attention with Negative Weights
di: Lv, Ang, et al.
Pubblicazione: (2024)
di: Lv, Ang, et al.
Pubblicazione: (2024)
Why Softmax Attention Outperforms Linear Attention
di: Deng, Yichuan, et al.
Pubblicazione: (2023)
di: Deng, Yichuan, et al.
Pubblicazione: (2023)
SEA: Sparse Linear Attention with Estimated Attention Mask
di: Lee, Heejun, et al.
Pubblicazione: (2023)
di: Lee, Heejun, et al.
Pubblicazione: (2023)
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
di: Fan, Zehao, et al.
Pubblicazione: (2025)
di: Fan, Zehao, et al.
Pubblicazione: (2025)
Unifying Linear-Time Attention via Latent Probabilistic Modelling
di: Dolga, Rares, et al.
Pubblicazione: (2024)
di: Dolga, Rares, et al.
Pubblicazione: (2024)
Linear Attention Sequence Parallelism
di: Sun, Weigao, et al.
Pubblicazione: (2024)
di: Sun, Weigao, et al.
Pubblicazione: (2024)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
di: Deng, Difan, et al.
Pubblicazione: (2026)
di: Deng, Difan, et al.
Pubblicazione: (2026)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
di: He, Mutian, et al.
Pubblicazione: (2025)
di: He, Mutian, et al.
Pubblicazione: (2025)
Scaling Linear Attention with Sparse State Expansion
di: Pan, Yuqi, et al.
Pubblicazione: (2025)
di: Pan, Yuqi, et al.
Pubblicazione: (2025)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
di: Van Nguyen, Chien, et al.
Pubblicazione: (2024)
di: Van Nguyen, Chien, et al.
Pubblicazione: (2024)
Learning Linear Attention in Polynomial Time
di: Yau, Morris, et al.
Pubblicazione: (2024)
di: Yau, Morris, et al.
Pubblicazione: (2024)
Higher-order Linear Attention
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
di: Yang, Songlin, et al.
Pubblicazione: (2023)
di: Yang, Songlin, et al.
Pubblicazione: (2023)
RoPE Attention Can Be Trained in Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
Training Tensor Attention Efficiently: From Cubic to Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
di: Zhao, Mingkuan, et al.
Pubblicazione: (2026)
di: Zhao, Mingkuan, et al.
Pubblicazione: (2026)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
di: Gerami, Armin, et al.
Pubblicazione: (2025)
di: Gerami, Armin, et al.
Pubblicazione: (2025)
LoLA: Low-Rank Linear Attention With Sparse Caching
di: McDermott, Luke, et al.
Pubblicazione: (2025)
di: McDermott, Luke, et al.
Pubblicazione: (2025)
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
di: Le, Dong, et al.
Pubblicazione: (2026)
di: Le, Dong, et al.
Pubblicazione: (2026)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
di: Zhou, Chenyu, et al.
Pubblicazione: (2026)
di: Zhou, Chenyu, et al.
Pubblicazione: (2026)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
di: Dong, Yihe, et al.
Pubblicazione: (2025)
di: Dong, Yihe, et al.
Pubblicazione: (2025)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
di: De, Soham, et al.
Pubblicazione: (2024)
di: De, Soham, et al.
Pubblicazione: (2024)
Parallax: Parameterized Local Linear Attention for Language Modeling
di: Zuo, Yifei, et al.
Pubblicazione: (2026)
di: Zuo, Yifei, et al.
Pubblicazione: (2026)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
di: Bombari, Simone, et al.
Pubblicazione: (2024)
di: Bombari, Simone, et al.
Pubblicazione: (2024)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
di: Wong, Liang Ze
Pubblicazione: (2025)
di: Wong, Liang Ze
Pubblicazione: (2025)
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
di: Goldstein, Daniel, et al.
Pubblicazione: (2025)
di: Goldstein, Daniel, et al.
Pubblicazione: (2025)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
di: Jin, Zehao, et al.
Pubblicazione: (2026)
di: Jin, Zehao, et al.
Pubblicazione: (2026)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
di: Qiu, Quantong, et al.
Pubblicazione: (2026)
di: Qiu, Quantong, et al.
Pubblicazione: (2026)
Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
di: Roy, Debjyoti Saha, et al.
Pubblicazione: (2024)
di: Roy, Debjyoti Saha, et al.
Pubblicazione: (2024)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
di: Chen, Lida, et al.
Pubblicazione: (2025)
di: Chen, Lida, et al.
Pubblicazione: (2025)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
di: Lee, Chungpa, et al.
Pubblicazione: (2026)
di: Lee, Chungpa, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Kimi Linear: An Expressive, Efficient Attention Architecture
di: Kimi Team, et al.
Pubblicazione: (2025) -
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
di: Zhang, Michael, et al.
Pubblicazione: (2024) -
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
di: Xiao, Kaicheng, et al.
Pubblicazione: (2026) -
GLU Attention Improve Transformer
di: Wang, Zehao
Pubblicazione: (2025) -
Routing Absorption in Sparse Attention: Why Random Gates Are Hard to Beat
di: Aquino-Michaels, Keston
Pubblicazione: (2026)