Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Weijie, Liu, Mingquan, Wang, Bolun, Wu, Simo, Xie, Nuobei, Zhu, Rui-Jie, Zhou, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
XicorAttention: Time Series Transformer Using Attention with Nonlinear Correlation
by: Kimura, Daichi, et al.
Published: (2025)
by: Kimura, Daichi, et al.
Published: (2025)
PRIME: Prototype-Driven Multimodal Pretraining for Cancer Prognosis with Missing Modalities
by: Yu, Kai, et al.
Published: (2026)
by: Yu, Kai, et al.
Published: (2026)
On Catastrophic Inheritance of Large Foundation Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Hybrid Focal and Full-Range Attention Based Graph Transformers
by: Zhu, Minhong, et al.
Published: (2023)
by: Zhu, Minhong, et al.
Published: (2023)
Attention Beyond Neighborhoods: Reviving Transformer for Graph Clustering
by: Xie, Xuanting, et al.
Published: (2025)
by: Xie, Xuanting, et al.
Published: (2025)
Scaling Reasoning without Attention
by: Zhao, Xueliang, et al.
Published: (2025)
by: Zhao, Xueliang, et al.
Published: (2025)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
by: Hu, Wenjie, et al.
Published: (2025)
by: Hu, Wenjie, et al.
Published: (2025)
TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
A Multi-Scale Graph Neural Process with Cross-Drug Co-Attention for Drug-Drug Interactions Prediction
by: Yan, Zimo, et al.
Published: (2025)
by: Yan, Zimo, et al.
Published: (2025)
InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model
by: Wang, Youjin, et al.
Published: (2026)
by: Wang, Youjin, et al.
Published: (2026)
GIAT: A Geologically-Informed Attention Transformer for Lithology Identification
by: Li, Jie, et al.
Published: (2026)
by: Li, Jie, et al.
Published: (2026)
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
Scaling Graph Transformers: A Comparative Study of Sparse and Dense Attention
by: Dimitrov, Leon
Published: (2025)
by: Dimitrov, Leon
Published: (2025)
Stable Attention Response for Reliable Precipitation Nowcasting
by: Wen, Penghui, et al.
Published: (2026)
by: Wen, Penghui, et al.
Published: (2026)
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
by: Bu, Rui, et al.
Published: (2025)
by: Bu, Rui, et al.
Published: (2025)
Summer-22B: A Systematic Approach to Dataset Engineering and Training at Scale for Video Foundation Model
by: Ryu, Simo, et al.
Published: (2026)
by: Ryu, Simo, et al.
Published: (2026)
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
by: Zhang, Yukun, et al.
Published: (2025)
by: Zhang, Yukun, et al.
Published: (2025)
Deep learning with noisy labels in medical prediction problems: a scoping review
by: Wei, Yishu, et al.
Published: (2024)
by: Wei, Yishu, et al.
Published: (2024)
How Transformers Learn to Plan via Multi-Token Prediction
by: Huang, Jianhao, et al.
Published: (2026)
by: Huang, Jianhao, et al.
Published: (2026)
Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
by: Saxena, Krati, et al.
Published: (2025)
by: Saxena, Krati, et al.
Published: (2025)
Exploring the Global-to-Local Attention Scheme in Graph Transformers: An Empirical Study
by: Wu, Gang, et al.
Published: (2025)
by: Wu, Gang, et al.
Published: (2025)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
by: Bae, Jeongin, et al.
Published: (2026)
by: Bae, Jeongin, et al.
Published: (2026)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
by: Li, Miaomiao, et al.
Published: (2025)
by: Li, Miaomiao, et al.
Published: (2025)
Unveiling and Controlling Anomalous Attention Distribution in Transformers
by: Yan, Ruiqing, et al.
Published: (2024)
by: Yan, Ruiqing, et al.
Published: (2024)
Embedding Reliability Verification Constraints into Generation Expansion Planning
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
Stem: Rethinking Causal Information Flow in Sparse Attention
by: Niu, Lin, et al.
Published: (2026)
by: Niu, Lin, et al.
Published: (2026)
Causally Sufficient and Necessary Feature Expansion for Class-Incremental Learning
by: Zhang, Zhen, et al.
Published: (2026)
by: Zhang, Zhen, et al.
Published: (2026)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
by: Zhou, Jingbo, et al.
Published: (2026)
by: Zhou, Jingbo, et al.
Published: (2026)
Beyond Linearity in Attention Projections: The Case for Nonlinear Queries
by: Karbevski, Marko
Published: (2026)
by: Karbevski, Marko
Published: (2026)
Multi-Item-Query Attention for Stable Sequential Recommendation
by: Xu, Mingshi, et al.
Published: (2025)
by: Xu, Mingshi, et al.
Published: (2025)
The Bayesian Geometry of Transformer Attention
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
Multi-Granular Attention based Heterogeneous Hypergraph Neural Network
by: Jin, Hong, et al.
Published: (2025)
by: Jin, Hong, et al.
Published: (2025)
MISE: Meta-knowledge Inheritance for Social Media-Based Stressor Estimation
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix
by: Zhang, Jinhao, et al.
Published: (2026)
by: Zhang, Jinhao, et al.
Published: (2026)
QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
by: Zhao, Zhixiong, et al.
Published: (2025)
by: Zhao, Zhixiong, et al.
Published: (2025)
Adjusting the Output of Decision Transformer with Action Gradient
by: Lin, Rui, et al.
Published: (2025)
by: Lin, Rui, et al.
Published: (2025)
MSCMHMST: A traffic flow prediction model based on Transformer
by: Geng, Weiyang, et al.
Published: (2025)
by: Geng, Weiyang, et al.
Published: (2025)
Data Fusion-Enhanced Decision Transformer for Stable Cross-Domain Generalization
by: Wang, Guojian, et al.
Published: (2025)
by: Wang, Guojian, et al.
Published: (2025)
Similar Items
-
XicorAttention: Time Series Transformer Using Attention with Nonlinear Correlation
by: Kimura, Daichi, et al.
Published: (2025) -
PRIME: Prototype-Driven Multimodal Pretraining for Cancer Prognosis with Missing Modalities
by: Yu, Kai, et al.
Published: (2026) -
On Catastrophic Inheritance of Large Foundation Models
by: Chen, Hao, et al.
Published: (2024) -
Hybrid Focal and Full-Range Attention Based Graph Transformers
by: Zhu, Minhong, et al.
Published: (2023) -
Attention Beyond Neighborhoods: Reviving Transformer for Graph Clustering
by: Xie, Xuanting, et al.
Published: (2025)