From TLinFormer to TConstFormer: The Leap to Constant-Time Transformer Attention: Achieving O(1) Computation and O(1) KV Cache during Autoregressive Inference
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Tang, Zhongpan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking Transformer Connectivity: TLinFormer, A Path to Exact, Full Context-Aware Linear Attention
von: Tang, Zhongpan
Veröffentlicht: (2025)
von: Tang, Zhongpan
Veröffentlicht: (2025)
CacheFormer: High Attention-Based Segment Caching
von: Singh, Sushant, et al.
Veröffentlicht: (2025)
von: Singh, Sushant, et al.
Veröffentlicht: (2025)
TimeFormer: Transformer with Attention Modulation Empowered by Temporal Characteristics for Time Series Forecasting
von: Liu, Zhipeng, et al.
Veröffentlicht: (2025)
von: Liu, Zhipeng, et al.
Veröffentlicht: (2025)
Compression is Routing: Reconstruction Error as an Intrinsic Signal for Modular Language Models
von: Tang, Zhongpan
Veröffentlicht: (2025)
von: Tang, Zhongpan
Veröffentlicht: (2025)
SFi-Former: Sparse Flow Induced Attention for Graph Transformer
von: Li, Zhonghao, et al.
Veröffentlicht: (2025)
von: Li, Zhonghao, et al.
Veröffentlicht: (2025)
NoiseFormer -- Noise Diffused Symmetric Attention Transformer
von: Kumar, Phani, et al.
Veröffentlicht: (2026)
von: Kumar, Phani, et al.
Veröffentlicht: (2026)
IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs
von: Mao, Yuzhen, et al.
Veröffentlicht: (2024)
von: Mao, Yuzhen, et al.
Veröffentlicht: (2024)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
MatFormer: Nested Transformer for Elastic Inference
von: Devvrit, et al.
Veröffentlicht: (2023)
von: Devvrit, et al.
Veröffentlicht: (2023)
The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference
von: Chodavarapu, Ranjith, et al.
Veröffentlicht: (2026)
von: Chodavarapu, Ranjith, et al.
Veröffentlicht: (2026)
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
Inference-Time Hyper-Scaling with KV Cache Compression
von: Łańcucki, Adrian, et al.
Veröffentlicht: (2025)
von: Łańcucki, Adrian, et al.
Veröffentlicht: (2025)
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
von: Shan, Jiquan, et al.
Veröffentlicht: (2025)
von: Shan, Jiquan, et al.
Veröffentlicht: (2025)
RecurFormer: Not All Transformer Heads Need Self-Attention
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
Breaking the KV Cache Bottleneck: Fan Duality Model Achieves O(1) Decode Memory with Superior Associative Recall
von: Fan, Yasong
Veröffentlicht: (2026)
von: Fan, Yasong
Veröffentlicht: (2026)
RTA-Former: Reverse Transformer Attention for Polyp Segmentation
von: Li, Zhikai, et al.
Veröffentlicht: (2024)
von: Li, Zhikai, et al.
Veröffentlicht: (2024)
BasisFormer: Attention-based Time Series Forecasting with Learnable and Interpretable Basis
von: Ni, Zelin, et al.
Veröffentlicht: (2023)
von: Ni, Zelin, et al.
Veröffentlicht: (2023)
Peri-midFormer: Periodic Pyramid Transformer for Time Series Analysis
von: Wu, Qiang, et al.
Veröffentlicht: (2024)
von: Wu, Qiang, et al.
Veröffentlicht: (2024)
ContiFormer: Continuous-Time Transformer for Irregular Time Series Modeling
von: Chen, Yuqi, et al.
Veröffentlicht: (2024)
von: Chen, Yuqi, et al.
Veröffentlicht: (2024)
ShapeFormer: Shapelet Transformer for Multivariate Time Series Classification
von: Le, Xuan-May, et al.
Veröffentlicht: (2024)
von: Le, Xuan-May, et al.
Veröffentlicht: (2024)
RiemannFormer: A Framework for Attention in Curved Spaces
von: Ji, Zhongping
Veröffentlicht: (2025)
von: Ji, Zhongping
Veröffentlicht: (2025)
Training Transformers for KV Cache Compressibility
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
von: Zhou, Jingbo, et al.
Veröffentlicht: (2026)
von: Zhou, Jingbo, et al.
Veröffentlicht: (2026)
ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
von: Zhang, Qiuyang, et al.
Veröffentlicht: (2026)
von: Zhang, Qiuyang, et al.
Veröffentlicht: (2026)
KV Cache Transform Coding for Compact Storage in LLM Inference
von: Staniszewski, Konrad, et al.
Veröffentlicht: (2025)
von: Staniszewski, Konrad, et al.
Veröffentlicht: (2025)
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel
von: Cutler, Dylan, et al.
Veröffentlicht: (2025)
von: Cutler, Dylan, et al.
Veröffentlicht: (2025)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
TeamFormer: Shallow Parallel Transformers with Progressive Approximation
von: Wang, Wei, et al.
Veröffentlicht: (2025)
von: Wang, Wei, et al.
Veröffentlicht: (2025)
WaveFormer: Wavelet Embedding Transformer for Biomedical Signals
von: Irani, Habib, et al.
Veröffentlicht: (2026)
von: Irani, Habib, et al.
Veröffentlicht: (2026)
From Complex Dynamics to DynFormer: Rethinking Transformers for PDEs
von: Lai, Pengyu, et al.
Veröffentlicht: (2026)
von: Lai, Pengyu, et al.
Veröffentlicht: (2026)
Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache
von: Dehghankar, Mohsen, et al.
Veröffentlicht: (2026)
von: Dehghankar, Mohsen, et al.
Veröffentlicht: (2026)
SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and $\mathcal{O}(T)$ Complexity
von: Zou, Shihao, et al.
Veröffentlicht: (2025)
von: Zou, Shihao, et al.
Veröffentlicht: (2025)
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
von: Li, Kunjun, et al.
Veröffentlicht: (2025)
von: Li, Kunjun, et al.
Veröffentlicht: (2025)
Compute Or Load KV Cache? Why Not Both?
von: Jin, Shuowei, et al.
Veröffentlicht: (2024)
von: Jin, Shuowei, et al.
Veröffentlicht: (2024)
RingFormer: A Neural Vocoder with Ring Attention and Convolution-Augmented Transformer
von: Hong, Seongho, et al.
Veröffentlicht: (2025)
von: Hong, Seongho, et al.
Veröffentlicht: (2025)
Decision ConvFormer: Local Filtering in MetaFormer is Sufficient for Decision Making
von: Kim, Jeonghye, et al.
Veröffentlicht: (2023)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2023)
JetFormer: An Autoregressive Generative Model of Raw Images and Text
von: Tschannen, Michael, et al.
Veröffentlicht: (2024)
von: Tschannen, Michael, et al.
Veröffentlicht: (2024)
TwinFormer: A Dual-Level Transformer for Long-Sequence Time-Series Forecasting
von: Kumavat, Mahima, et al.
Veröffentlicht: (2025)
von: Kumavat, Mahima, et al.
Veröffentlicht: (2025)
The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026)
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Rethinking Transformer Connectivity: TLinFormer, A Path to Exact, Full Context-Aware Linear Attention
von: Tang, Zhongpan
Veröffentlicht: (2025) -
CacheFormer: High Attention-Based Segment Caching
von: Singh, Sushant, et al.
Veröffentlicht: (2025) -
TimeFormer: Transformer with Attention Modulation Empowered by Temporal Characteristics for Time Series Forecasting
von: Liu, Zhipeng, et al.
Veröffentlicht: (2025) -
Compression is Routing: Reconstruction Error as an Intrinsic Signal for Modular Language Models
von: Tang, Zhongpan
Veröffentlicht: (2025) -
SFi-Former: Sparse Flow Induced Attention for Graph Transformer
von: Li, Zhonghao, et al.
Veröffentlicht: (2025)