SchoenbAt: Rethinking Attention with Polynomial basis
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Yuhan, Ding, Lizhong, Yang, Yuwan, Guo, Xuewei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Macformer: Transformer with Random Maclaurin Feature Attention
by: Guo, Yuhan, et al.
Published: (2024)
by: Guo, Yuhan, et al.
Published: (2024)
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
by: Guo, Yuhan, et al.
Published: (2025)
by: Guo, Yuhan, et al.
Published: (2025)
AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
MatrixKAN: Parallelized Kolmogorov-Arnold Network
by: Coffman, Cale, et al.
Published: (2025)
by: Coffman, Cale, et al.
Published: (2025)
Dependence Induced Representations
by: Xu, Xiangxiang, et al.
Published: (2024)
by: Xu, Xiangxiang, et al.
Published: (2024)
Neural Feature Learning in Function Space
by: Xu, Xiangxiang, et al.
Published: (2023)
by: Xu, Xiangxiang, et al.
Published: (2023)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
by: Ding, Yifu, et al.
Published: (2026)
by: Ding, Yifu, et al.
Published: (2026)
Beyond Classical Attention: Quantum Attention for Scalable Computation
by: Guo, Xuyang, et al.
Published: (2023)
by: Guo, Xuyang, et al.
Published: (2023)
CGCMA: Conditionally-Gated Cross-Modal Attention for Event-Conditioned Asynchronous Fusion
by: Guo, Yunxiang
Published: (2026)
by: Guo, Yunxiang
Published: (2026)
Separable Computation of Information Measures
by: Xu, Xiangxiang, et al.
Published: (2025)
by: Xu, Xiangxiang, et al.
Published: (2025)
Scaling Context Requires Rethinking Attention
by: Gelada, Carles, et al.
Published: (2025)
by: Gelada, Carles, et al.
Published: (2025)
QuantKAN: A Unified Quantization Framework for Kolmogorov Arnold Networks
by: Fuad, Kazi Ahmed Asif, et al.
Published: (2025)
by: Fuad, Kazi Ahmed Asif, et al.
Published: (2025)
Why does Knowledge Distillation Work? Rethink its Attention and Fidelity Mechanism
by: Guo, Chenqi, et al.
Published: (2024)
by: Guo, Chenqi, et al.
Published: (2024)
Log-Linear Attention
by: Guo, Han, et al.
Published: (2025)
by: Guo, Han, et al.
Published: (2025)
Unveiling and Causalizing CoT: A Causal Pespective
by: Fu, Jiarun, et al.
Published: (2025)
by: Fu, Jiarun, et al.
Published: (2025)
Rethinking the Graph Polynomial Filter via Positive and Negative Coupling Analysis
by: Wen, Haodong, et al.
Published: (2024)
by: Wen, Haodong, et al.
Published: (2024)
Rethinking Irregular Time Series Forecasting: A Simple yet Effective Baseline
by: Liu, Xvyuan, et al.
Published: (2025)
by: Liu, Xvyuan, et al.
Published: (2025)
Shallow Neural Networks Learn Low-Degree Spherical Polynomials with Feature Learning by Learnable Channel Attention
by: Yang, Yingzhen
Published: (2025)
by: Yang, Yingzhen
Published: (2025)
Rethinking Multi-Modal Learning from Gradient Uncertainty
by: Guo, Peizheng, et al.
Published: (2025)
by: Guo, Peizheng, et al.
Published: (2025)
FlashKAT: Understanding and Addressing Performance Bottlenecks in the Kolmogorov-Arnold Transformer
by: Raffel, Matthew, et al.
Published: (2025)
by: Raffel, Matthew, et al.
Published: (2025)
You Need Better Attention Priors
by: Litman, Elon, et al.
Published: (2026)
by: Litman, Elon, et al.
Published: (2026)
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
by: Vaish, Puru, et al.
Published: (2024)
by: Vaish, Puru, et al.
Published: (2024)
Greedy-Gnorm: A Gradient Matrix Norm-Based Alternative to Attention Entropy for Head Pruning
by: Guo, Yuxi, et al.
Published: (2026)
by: Guo, Yuxi, et al.
Published: (2026)
Stem: Rethinking Causal Information Flow in Sparse Attention
by: Niu, Lin, et al.
Published: (2026)
by: Niu, Lin, et al.
Published: (2026)
Decentralized Attention Fails Centralized Signals: Rethinking Transformers for Medical Time Series
by: Yu, Guoqi, et al.
Published: (2026)
by: Yu, Guoqi, et al.
Published: (2026)
Heterophily-Aware Graph Attention Network
by: Wang, Junfu, et al.
Published: (2023)
by: Wang, Junfu, et al.
Published: (2023)
Fast KV Compaction via Attention Matching
by: Zweiger, Adam, et al.
Published: (2026)
by: Zweiger, Adam, et al.
Published: (2026)
BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System
by: Raffel, Matthew, et al.
Published: (2025)
by: Raffel, Matthew, et al.
Published: (2025)
Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous Translation
by: Raffel, Matthew, et al.
Published: (2024)
by: Raffel, Matthew, et al.
Published: (2024)
LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned Proportions
by: Agostinelli, Victor, et al.
Published: (2024)
by: Agostinelli, Victor, et al.
Published: (2024)
FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction
by: Kulp, Gabriel, et al.
Published: (2024)
by: Kulp, Gabriel, et al.
Published: (2024)
Rethinking Link Prediction for Directed Graphs
by: He, Mingguo, et al.
Published: (2025)
by: He, Mingguo, et al.
Published: (2025)
Generalization and Risk Bounds for Recurrent Neural Networks
by: Cheng, Xuewei, et al.
Published: (2024)
by: Cheng, Xuewei, et al.
Published: (2024)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
Multi-Head Low-Rank Attention
by: Liu, Songtao, et al.
Published: (2026)
by: Liu, Songtao, et al.
Published: (2026)
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
by: Ding, Bowen, et al.
Published: (2025)
by: Ding, Bowen, et al.
Published: (2025)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
by: Yao, Xinhao, et al.
Published: (2025)
by: Yao, Xinhao, et al.
Published: (2025)
Learning Linear Attention in Polynomial Time
by: Yau, Morris, et al.
Published: (2024)
by: Yau, Morris, et al.
Published: (2024)
Similar Items
-
Macformer: Transformer with Random Maclaurin Feature Attention
by: Guo, Yuhan, et al.
Published: (2024) -
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024) -
DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
by: Guo, Yuhan, et al.
Published: (2025) -
AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
by: Zhang, Yu, et al.
Published: (2025) -
MatrixKAN: Parallelized Kolmogorov-Arnold Network
by: Coffman, Cale, et al.
Published: (2025)