Unifying Linear-Time Attention via Latent Probabilistic Modelling
Fuente:
arXiv
Saved in:
| Main Authors: | Dolga, Rares, Maystre, Lucas, Cobzarenco, Marius, Barber, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
by: Dolga, Rares, et al.
Published: (2025)
by: Dolga, Rares, et al.
Published: (2025)
RotRNN: Modelling Long Sequences with Rotations
by: Biegun, Kai, et al.
Published: (2024)
by: Biegun, Kai, et al.
Published: (2024)
Incremental Sequence Classification with Temporal Consistency
by: Maystre, Lucas, et al.
Published: (2025)
by: Maystre, Lucas, et al.
Published: (2025)
When Embedding Models Meet: Procrustes Bounds and Applications
by: Maystre, Lucas, et al.
Published: (2025)
by: Maystre, Lucas, et al.
Published: (2025)
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
by: Deng, Difan, et al.
Published: (2026)
by: Deng, Difan, et al.
Published: (2026)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
by: Rizvi-Martel, Michael, et al.
Published: (2026)
by: Rizvi-Martel, Michael, et al.
Published: (2026)
Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
by: Jin, Zehao, et al.
Published: (2026)
by: Jin, Zehao, et al.
Published: (2026)
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
by: Tang, Xiaohang, et al.
Published: (2025)
by: Tang, Xiaohang, et al.
Published: (2025)
Transformer-VQ: Linear-Time Transformers via Vector Quantization
by: Lingle, Lucas D.
Published: (2023)
by: Lingle, Lucas D.
Published: (2023)
Generalized Probabilistic Attention Mechanism in Transformers
by: Heo, DongNyeong, et al.
Published: (2024)
by: Heo, DongNyeong, et al.
Published: (2024)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
SEA: Sparse Linear Attention with Estimated Attention Mask
by: Lee, Heejun, et al.
Published: (2023)
by: Lee, Heejun, et al.
Published: (2023)
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
by: De, Soham, et al.
Published: (2024)
by: De, Soham, et al.
Published: (2024)
Latent Distribution Decoupling: A Probabilistic Framework for Uncertainty-Aware Multimodal Emotion Recognition
by: Huang, Jingwang, et al.
Published: (2025)
by: Huang, Jingwang, et al.
Published: (2025)
Latent Thought Models with Variational Bayes Inference-Time Computation
by: Kong, Deqian, et al.
Published: (2025)
by: Kong, Deqian, et al.
Published: (2025)
Towards Logically Consistent Language Models via Probabilistic Reasoning
by: Calanzone, Diego, et al.
Published: (2024)
by: Calanzone, Diego, et al.
Published: (2024)
Scaling Linear Attention with Sparse State Expansion
by: Pan, Yuqi, et al.
Published: (2025)
by: Pan, Yuqi, et al.
Published: (2025)
Kimi Linear: An Expressive, Efficient Attention Architecture
by: Kimi Team, et al.
Published: (2025)
by: Kimi Team, et al.
Published: (2025)
Training Tensor Attention Efficiently: From Cubic to Almost Linear Time
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
RoPE Attention Can Be Trained in Almost Linear Time
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
by: He, Mutian, et al.
Published: (2025)
by: He, Mutian, et al.
Published: (2025)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
by: Dentamaro, Vincenzo
Published: (2025)
by: Dentamaro, Vincenzo
Published: (2025)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
by: Zhang, Michael, et al.
Published: (2024)
by: Zhang, Michael, et al.
Published: (2024)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Learning Linear Attention in Polynomial Time
by: Yau, Morris, et al.
Published: (2024)
by: Yau, Morris, et al.
Published: (2024)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
by: Lee, Chungpa, et al.
Published: (2026)
by: Lee, Chungpa, et al.
Published: (2026)
Parallax: Parameterized Local Linear Attention for Language Modeling
by: Zuo, Yifei, et al.
Published: (2026)
by: Zuo, Yifei, et al.
Published: (2026)
Inference-Time Rethinking with Latent Thought Vectors for Math Reasoning
by: Kong, Deqian, et al.
Published: (2026)
by: Kong, Deqian, et al.
Published: (2026)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
by: Xiao, Kaicheng, et al.
Published: (2026)
by: Xiao, Kaicheng, et al.
Published: (2026)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
by: Gerami, Armin, et al.
Published: (2025)
by: Gerami, Armin, et al.
Published: (2025)
LoLA: Low-Rank Linear Attention With Sparse Caching
by: McDermott, Luke, et al.
Published: (2025)
by: McDermott, Luke, et al.
Published: (2025)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
by: Knupp, Jonas, et al.
Published: (2026)
by: Knupp, Jonas, et al.
Published: (2026)
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
by: Rohweder, Jonas, et al.
Published: (2026)
by: Rohweder, Jonas, et al.
Published: (2026)
Higher-order Linear Attention
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
ChatTime: A Unified Multimodal Time Series Foundation Model Bridging Numerical and Textual Data
by: Wang, Chengsen, et al.
Published: (2024)
by: Wang, Chengsen, et al.
Published: (2024)
Probabilistic Topic Modelling with Transformer Representations
by: Reuter, Arik, et al.
Published: (2024)
by: Reuter, Arik, et al.
Published: (2024)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
by: Fu, Zichuan, et al.
Published: (2026)
by: Fu, Zichuan, et al.
Published: (2026)
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
by: Le, Dong, et al.
Published: (2026)
by: Le, Dong, et al.
Published: (2026)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
by: Zhou, Chenyu, et al.
Published: (2026)
by: Zhou, Chenyu, et al.
Published: (2026)
Similar Items
-
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
by: Dolga, Rares, et al.
Published: (2025) -
RotRNN: Modelling Long Sequences with Rotations
by: Biegun, Kai, et al.
Published: (2024) -
Incremental Sequence Classification with Temporal Consistency
by: Maystre, Lucas, et al.
Published: (2025) -
When Embedding Models Meet: Procrustes Bounds and Applications
by: Maystre, Lucas, et al.
Published: (2025) -
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
by: Deng, Difan, et al.
Published: (2026)