Cottention: Linear Transformers With Cosine Attention
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mongaras, Gabriel, Dohm, Trevor, Larson, Eric C. |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
par: Mongaras, Gabriel, et autres
Publié: (2025)
par: Mongaras, Gabriel, et autres
Publié: (2025)
2Mamba2Furious: Linear in Complexity, Competitive in Accuracy
par: Mongaras, Gabriel, et autres
Publié: (2026)
par: Mongaras, Gabriel, et autres
Publié: (2026)
Discrete Cosine Transform Based Decorrelated Attention for Vision Transformers
par: Pan, Hongyi, et autres
Publié: (2024)
par: Pan, Hongyi, et autres
Publié: (2024)
Efficient Neural Networks with Discrete Cosine Transform Activations
par: Martinez-Gost, Marc, et autres
Publié: (2025)
par: Martinez-Gost, Marc, et autres
Publié: (2025)
Adaptive function approximation based on the Discrete Cosine Transform (DCT)
par: Pérez-Neira, Ana I., et autres
Publié: (2023)
par: Pérez-Neira, Ana I., et autres
Publié: (2023)
Log-Linear Attention
par: Guo, Han, et autres
Publié: (2025)
par: Guo, Han, et autres
Publié: (2025)
InAttention: Linear Context Scaling for Transformers
par: Eisner, Joseph
Publié: (2024)
par: Eisner, Joseph
Publié: (2024)
The Hidden Pitfalls of the Cosine Similarity Loss
par: Draganov, Andrew, et autres
Publié: (2024)
par: Draganov, Andrew, et autres
Publié: (2024)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
par: Hu, Wenjie, et autres
Publié: (2025)
par: Hu, Wenjie, et autres
Publié: (2025)
Variance-Adjusted Cosine Distance as Similarity Metric
par: Sahoo, Satyajeet, et autres
Publié: (2025)
par: Sahoo, Satyajeet, et autres
Publié: (2025)
On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
par: Ro, Yeonju, et autres
Publié: (2025)
par: Ro, Yeonju, et autres
Publié: (2025)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
par: Nguyen, Huy, et autres
Publié: (2024)
par: Nguyen, Huy, et autres
Publié: (2024)
Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers
par: Chen, Brian K, et autres
Publié: (2024)
par: Chen, Brian K, et autres
Publié: (2024)
In Defense of Cosine Similarity: Normalization Eliminates the Gauge Freedom
par: Bouhsine, Taha
Publié: (2026)
par: Bouhsine, Taha
Publié: (2026)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
par: Xie, Zixuan, et autres
Publié: (2026)
par: Xie, Zixuan, et autres
Publié: (2026)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
par: Pandey, Vishal, et autres
Publié: (2026)
par: Pandey, Vishal, et autres
Publié: (2026)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
par: Lu, Jiecheng, et autres
Publié: (2026)
par: Lu, Jiecheng, et autres
Publié: (2026)
Gated Linear Attention Transformers with Hardware-Efficient Training
par: Yang, Songlin, et autres
Publié: (2023)
par: Yang, Songlin, et autres
Publié: (2023)
CosineGate: Semantic Dynamic Routing via Cosine Incompatibility in Residual Networks
par: Thota, Yogeswar Reddy
Publié: (2025)
par: Thota, Yogeswar Reddy
Publié: (2025)
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
par: Wu, Ziyang, et autres
Publié: (2024)
par: Wu, Ziyang, et autres
Publié: (2024)
Semantics at an Angle: When Cosine Similarity Works Until It Doesn't
par: You, Kisung
Publié: (2025)
par: You, Kisung
Publié: (2025)
Gamma Mixture Modeling for Cosine Similarity in Small Language Models
par: Player, Kevin
Publié: (2025)
par: Player, Kevin
Publié: (2025)
Is Cosine-Similarity of Embeddings Really About Similarity?
par: Steck, Harald, et autres
Publié: (2024)
par: Steck, Harald, et autres
Publié: (2024)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
par: Gerami, Armin, et autres
Publié: (2025)
par: Gerami, Armin, et autres
Publié: (2025)
The Cosine Schedule is Fisher-Rao-Optimal for Masked Discrete Diffusion Models
par: Zhang, Leo, et autres
Publié: (2025)
par: Zhang, Leo, et autres
Publié: (2025)
Outlier Detection Using Vector Cosine Similarity by Adding a Dimension
par: Shen, Zhongyang
Publié: (2025)
par: Shen, Zhongyang
Publié: (2025)
EDiT: Efficient Diffusion Transformers with Linear Compressed Attention
par: Becker, Philipp, et autres
Publié: (2025)
par: Becker, Philipp, et autres
Publié: (2025)
Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
par: Wi, Hyowon, et autres
Publié: (2025)
par: Wi, Hyowon, et autres
Publié: (2025)
SAVGO: Learning State-Action Value Geometry with Cosine Similarity for Continuous Control
par: Orfanoudakis, Stavros, et autres
Publié: (2026)
par: Orfanoudakis, Stavros, et autres
Publié: (2026)
Cosine Scoring with Uncertainty for Neural Speaker Embedding
par: Wang, Qiongqiong, et autres
Publié: (2024)
par: Wang, Qiongqiong, et autres
Publié: (2024)
Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive Forecasting
par: Lu, Jiecheng, et autres
Publié: (2025)
par: Lu, Jiecheng, et autres
Publié: (2025)
Multi-Margin Cosine Loss: Proposal and Application in Recommender Systems
par: Ozsoy, Makbule Gulcin
Publié: (2024)
par: Ozsoy, Makbule Gulcin
Publié: (2024)
Rethinking Transformer Connectivity: TLinFormer, A Path to Exact, Full Context-Aware Linear Attention
par: Tang, Zhongpan
Publié: (2025)
par: Tang, Zhongpan
Publié: (2025)
TabFlex: Scaling Tabular Learning to Millions with Linear Attention
par: Zeng, Yuchen, et autres
Publié: (2025)
par: Zeng, Yuchen, et autres
Publié: (2025)
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
par: Singh, Vaibhav, et autres
Publié: (2025)
par: Singh, Vaibhav, et autres
Publié: (2025)
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
par: Goldstein, Daniel, et autres
Publié: (2025)
par: Goldstein, Daniel, et autres
Publié: (2025)
Exact Linear Attention
par: Ou, Weinuo
Publié: (2026)
par: Ou, Weinuo
Publié: (2026)
Kaczmarz Linear Attention
par: Zou, Jiaxuan, et autres
Publié: (2026)
par: Zou, Jiaxuan, et autres
Publié: (2026)
Scaling Continuous Kernels with Sparse Fourier Domain Learning
par: Harper, Clayton, et autres
Publié: (2024)
par: Harper, Clayton, et autres
Publié: (2024)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
par: Chen, Xingwu, et autres
Publié: (2024)
par: Chen, Xingwu, et autres
Publié: (2024)
Documents similaires
-
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
par: Mongaras, Gabriel, et autres
Publié: (2025) -
2Mamba2Furious: Linear in Complexity, Competitive in Accuracy
par: Mongaras, Gabriel, et autres
Publié: (2026) -
Discrete Cosine Transform Based Decorrelated Attention for Vision Transformers
par: Pan, Hongyi, et autres
Publié: (2024) -
Efficient Neural Networks with Discrete Cosine Transform Activations
par: Martinez-Gost, Marc, et autres
Publié: (2025) -
Adaptive function approximation based on the Discrete Cosine Transform (DCT)
par: Pérez-Neira, Ana I., et autres
Publié: (2023)