Rethinking Transformer Connectivity: TLinFormer, A Path to Exact, Full Context-Aware Linear Attention
Fuente:
arXiv
Guardado en:
| Autor principal: | Tang, Zhongpan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From TLinFormer to TConstFormer: The Leap to Constant-Time Transformer Attention: Achieving O(1) Computation and O(1) KV Cache during Autoregressive Inference
por: Tang, Zhongpan
Publicado: (2025)
por: Tang, Zhongpan
Publicado: (2025)
Compression is Routing: Reconstruction Error as an Intrinsic Signal for Modular Language Models
por: Tang, Zhongpan
Publicado: (2025)
por: Tang, Zhongpan
Publicado: (2025)
Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers
por: Chen, Brian K, et al.
Publicado: (2024)
por: Chen, Brian K, et al.
Publicado: (2024)
InAttention: Linear Context Scaling for Transformers
por: Eisner, Joseph
Publicado: (2024)
por: Eisner, Joseph
Publicado: (2024)
Exact Linear Attention
por: Ou, Weinuo
Publicado: (2026)
por: Ou, Weinuo
Publicado: (2026)
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
por: Wang, Haiyang, et al.
Publicado: (2024)
por: Wang, Haiyang, et al.
Publicado: (2024)
RingFormer: Rethinking Recurrent Transformer with Adaptive Level Signals
por: Heo, Jaemu, et al.
Publicado: (2025)
por: Heo, Jaemu, et al.
Publicado: (2025)
Exact Learning Dynamics of In-Context Learning in Linear Transformers and Its Application to Non-Linear Transformers
por: Mainali, Nischal, et al.
Publicado: (2025)
por: Mainali, Nischal, et al.
Publicado: (2025)
SFi-Former: Sparse Flow Induced Attention for Graph Transformer
por: Li, Zhonghao, et al.
Publicado: (2025)
por: Li, Zhonghao, et al.
Publicado: (2025)
TimeFormer: Transformer with Attention Modulation Empowered by Temporal Characteristics for Time Series Forecasting
por: Liu, Zhipeng, et al.
Publicado: (2025)
por: Liu, Zhipeng, et al.
Publicado: (2025)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
por: Xie, Zixuan, et al.
Publicado: (2026)
por: Xie, Zixuan, et al.
Publicado: (2026)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
por: Pandey, Vishal, et al.
Publicado: (2026)
por: Pandey, Vishal, et al.
Publicado: (2026)
Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
por: Lei, Jingdi, et al.
Publicado: (2025)
por: Lei, Jingdi, et al.
Publicado: (2025)
NoiseFormer -- Noise Diffused Symmetric Attention Transformer
por: Kumar, Phani, et al.
Publicado: (2026)
por: Kumar, Phani, et al.
Publicado: (2026)
Scaling Context Requires Rethinking Attention
por: Gelada, Carles, et al.
Publicado: (2025)
por: Gelada, Carles, et al.
Publicado: (2025)
From Complex Dynamics to DynFormer: Rethinking Transformers for PDEs
por: Lai, Pengyu, et al.
Publicado: (2026)
por: Lai, Pengyu, et al.
Publicado: (2026)
Power-based Partial Attention: Bridging Linear-Complexity and Full Attention
por: Huang, Yufeng
Publicado: (2026)
por: Huang, Yufeng
Publicado: (2026)
ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers
por: Hsu, Chih-Chung, et al.
Publicado: (2026)
por: Hsu, Chih-Chung, et al.
Publicado: (2026)
Superiority of Multi-Head Attention in In-Context Linear Regression
por: Cui, Yingqian, et al.
Publicado: (2024)
por: Cui, Yingqian, et al.
Publicado: (2024)
Generalized Linear Mode Connectivity for Transformers
por: Theus, Alexander, et al.
Publicado: (2025)
por: Theus, Alexander, et al.
Publicado: (2025)
Replacing Paths with Connection-Biased Attention for Knowledge Graph Completion
por: Dutta, Sharmishtha, et al.
Publicado: (2024)
por: Dutta, Sharmishtha, et al.
Publicado: (2024)
Training Dynamics of In-Context Learning in Linear Attention
por: Zhang, Yedi, et al.
Publicado: (2025)
por: Zhang, Yedi, et al.
Publicado: (2025)
Hybrid Dual-Path Linear Transformations for Efficient Transformer Architectures
por: Khasia, Vladimer
Publicado: (2026)
por: Khasia, Vladimer
Publicado: (2026)
Exact Attention Sensitivity and the Geometry of Transformer Stability
por: Emadi, Seyed Morteza
Publicado: (2026)
por: Emadi, Seyed Morteza
Publicado: (2026)
Can Transformers Learn Full Bayesian Inference in Context?
por: Reuter, Arik, et al.
Publicado: (2025)
por: Reuter, Arik, et al.
Publicado: (2025)
Cottention: Linear Transformers With Cosine Attention
por: Mongaras, Gabriel, et al.
Publicado: (2024)
por: Mongaras, Gabriel, et al.
Publicado: (2024)
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
por: Shan, Jiquan, et al.
Publicado: (2025)
por: Shan, Jiquan, et al.
Publicado: (2025)
RecurFormer: Not All Transformer Heads Need Self-Attention
por: Yan, Ruiqing, et al.
Publicado: (2024)
por: Yan, Ruiqing, et al.
Publicado: (2024)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
por: Chen, Xingwu, et al.
Publicado: (2024)
por: Chen, Xingwu, et al.
Publicado: (2024)
E2Former-V2: On-the-Fly Equivariant Attention with Linear Activation Memory
por: Huang, Lin, et al.
Publicado: (2026)
por: Huang, Lin, et al.
Publicado: (2026)
Linear Transformers are Versatile In-Context Learners
por: Vladymyrov, Max, et al.
Publicado: (2024)
por: Vladymyrov, Max, et al.
Publicado: (2024)
RiemannFormer: A Framework for Attention in Curved Spaces
por: Ji, Zhongping
Publicado: (2025)
por: Ji, Zhongping
Publicado: (2025)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
por: Zhou, Jingbo, et al.
Publicado: (2026)
por: Zhou, Jingbo, et al.
Publicado: (2026)
LinFormer: A Linear-based Lightweight Transformer Architecture For Time-Aware MIMO Channel Prediction
por: Jin, Yanliang, et al.
Publicado: (2024)
por: Jin, Yanliang, et al.
Publicado: (2024)
DeepCrossAttention: Supercharging Transformer Residual Connections
por: Heddes, Mike, et al.
Publicado: (2025)
por: Heddes, Mike, et al.
Publicado: (2025)
E2Former: An Efficient and Equivariant Transformer with Linear-Scaling Tensor Products
por: Li, Yunyang, et al.
Publicado: (2025)
por: Li, Yunyang, et al.
Publicado: (2025)
FieldFormer: Locality-Aware Transformers for Spatio-Temporal Modeling on Sparse Sensor Networks
por: Bhardwaj, Ankit, et al.
Publicado: (2025)
por: Bhardwaj, Ankit, et al.
Publicado: (2025)
RingFormer: A Neural Vocoder with Ring Attention and Convolution-Augmented Transformer
por: Hong, Seongho, et al.
Publicado: (2025)
por: Hong, Seongho, et al.
Publicado: (2025)
RTA-Former: Reverse Transformer Attention for Polyp Segmentation
por: Li, Zhikai, et al.
Publicado: (2024)
por: Li, Zhikai, et al.
Publicado: (2024)
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
por: Aggarwal, Shubham, et al.
Publicado: (2026)
por: Aggarwal, Shubham, et al.
Publicado: (2026)
Ejemplares similares
-
From TLinFormer to TConstFormer: The Leap to Constant-Time Transformer Attention: Achieving O(1) Computation and O(1) KV Cache during Autoregressive Inference
por: Tang, Zhongpan
Publicado: (2025) -
Compression is Routing: Reconstruction Error as an Intrinsic Signal for Modular Language Models
por: Tang, Zhongpan
Publicado: (2025) -
Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers
por: Chen, Brian K, et al.
Publicado: (2024) -
InAttention: Linear Context Scaling for Transformers
por: Eisner, Joseph
Publicado: (2024) -
Exact Linear Attention
por: Ou, Weinuo
Publicado: (2026)