Guardado en:
| Autores principales: | Li, Jie, Yang, Qishun, Li, Nuo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.09165 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
por: Adhikari, Rabin
Publicado: (2025)
por: Adhikari, Rabin
Publicado: (2025)
Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling
por: Zhao, Weijie, et al.
Publicado: (2026)
por: Zhao, Weijie, et al.
Publicado: (2026)
WLFM: A Well-Logs Foundation Model for Multi-Task and Cross-Well Geological Interpretation
por: Qi, Zhenyu, et al.
Publicado: (2025)
por: Qi, Zhenyu, et al.
Publicado: (2025)
Neurocircuitry-Inspired Hierarchical Graph Causal Attention Networks for Explainable Depression Identification
por: Chen, Weidao, et al.
Publicado: (2025)
por: Chen, Weidao, et al.
Publicado: (2025)
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
por: Bu, Rui, et al.
Publicado: (2025)
por: Bu, Rui, et al.
Publicado: (2025)
Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention
por: Freytes, Luis Rosario
Publicado: (2026)
por: Freytes, Luis Rosario
Publicado: (2026)
The Bayesian Geometry of Transformer Attention
por: Agarwal, Naman, et al.
Publicado: (2025)
por: Agarwal, Naman, et al.
Publicado: (2025)
Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
por: Saxena, Krati, et al.
Publicado: (2025)
por: Saxena, Krati, et al.
Publicado: (2025)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
por: Lu, Jiecheng, et al.
Publicado: (2026)
por: Lu, Jiecheng, et al.
Publicado: (2026)
Contrasformer: A Brain Network Contrastive Transformer for Neurodegenerative Condition Identification
por: Xu, Jiaxing, et al.
Publicado: (2024)
por: Xu, Jiaxing, et al.
Publicado: (2024)
What Matters in Transformers? Not All Attention is Needed
por: He, Shwai, et al.
Publicado: (2024)
por: He, Shwai, et al.
Publicado: (2024)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
por: Li, Cheng, et al.
Publicado: (2025)
por: Li, Cheng, et al.
Publicado: (2025)
CVTGAD: Simplified Transformer with Cross-View Attention for Unsupervised Graph-level Anomaly Detection
por: Li, Jindong, et al.
Publicado: (2024)
por: Li, Jindong, et al.
Publicado: (2024)
Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive Forecasting
por: Lu, Jiecheng, et al.
Publicado: (2025)
por: Lu, Jiecheng, et al.
Publicado: (2025)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
por: Zhou, Jingbo, et al.
Publicado: (2026)
por: Zhou, Jingbo, et al.
Publicado: (2026)
VSFormer: Value and Shape-Aware Transformer with Prior-Enhanced Self-Attention for Multivariate Time Series Classification
por: Xi, Wenjie, et al.
Publicado: (2024)
por: Xi, Wenjie, et al.
Publicado: (2024)
XicorAttention: Time Series Transformer Using Attention with Nonlinear Correlation
por: Kimura, Daichi, et al.
Publicado: (2025)
por: Kimura, Daichi, et al.
Publicado: (2025)
Synthetic Geology: Structural Geology Meets Deep Learning
por: Ghyselincks, Simon, et al.
Publicado: (2025)
por: Ghyselincks, Simon, et al.
Publicado: (2025)
A Multi-Scale Graph Neural Process with Cross-Drug Co-Attention for Drug-Drug Interactions Prediction
por: Yan, Zimo, et al.
Publicado: (2025)
por: Yan, Zimo, et al.
Publicado: (2025)
Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
por: Peng, Miao, et al.
Publicado: (2025)
por: Peng, Miao, et al.
Publicado: (2025)
ControlMath: Controllable Data Generation Promotes Math Generalist Models
por: Chen, Nuo, et al.
Publicado: (2024)
por: Chen, Nuo, et al.
Publicado: (2024)
UMoE: Unifying Attention and FFN with Shared Experts
por: Yang, Yuanhang, et al.
Publicado: (2025)
por: Yang, Yuanhang, et al.
Publicado: (2025)
A Graph Transformer-Driven Approach for Network Robustness Learning
por: Zhang, Yu, et al.
Publicado: (2023)
por: Zhang, Yu, et al.
Publicado: (2023)
CrowdTransfer: Enabling Crowd Knowledge Transfer in AIoT Community
por: Liu, Yan, et al.
Publicado: (2024)
por: Liu, Yan, et al.
Publicado: (2024)
PMET: Precise Model Editing in a Transformer
por: Li, Xiaopeng, et al.
Publicado: (2023)
por: Li, Xiaopeng, et al.
Publicado: (2023)
Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
por: Dhayalkar, Sahil Rajesh
Publicado: (2025)
por: Dhayalkar, Sahil Rajesh
Publicado: (2025)
Exact Attention Sensitivity and the Geometry of Transformer Stability
por: Emadi, Seyed Morteza
Publicado: (2026)
por: Emadi, Seyed Morteza
Publicado: (2026)
Unveiling and Controlling Anomalous Attention Distribution in Transformers
por: Yan, Ruiqing, et al.
Publicado: (2024)
por: Yan, Ruiqing, et al.
Publicado: (2024)
Graph Convolutions Enrich the Self-Attention in Transformers!
por: Choi, Jeongwhan, et al.
Publicado: (2023)
por: Choi, Jeongwhan, et al.
Publicado: (2023)
Higher-Order Transformers With Kronecker-Structured Attention
por: Omranpour, Soroush, et al.
Publicado: (2024)
por: Omranpour, Soroush, et al.
Publicado: (2024)
Expanding Expressivity in Transformer Models with MöbiusAttention
por: Halacheva, Anna-Maria, et al.
Publicado: (2024)
por: Halacheva, Anna-Maria, et al.
Publicado: (2024)
CAPS: Unifying Attention, Recurrence, and Alignment in Transformer-based Time Series Forecasting
por: Pati, Viresh, et al.
Publicado: (2026)
por: Pati, Viresh, et al.
Publicado: (2026)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
por: Hu, Wenjie, et al.
Publicado: (2025)
por: Hu, Wenjie, et al.
Publicado: (2025)
Decentralized Attention Fails Centralized Signals: Rethinking Transformers for Medical Time Series
por: Yu, Guoqi, et al.
Publicado: (2026)
por: Yu, Guoqi, et al.
Publicado: (2026)
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
por: Ke, Yekun, et al.
Publicado: (2024)
por: Ke, Yekun, et al.
Publicado: (2024)
AttentionSmithy: A Modular Framework for Rapid Transformer Development and Customization
por: Cranney, Caleb, et al.
Publicado: (2025)
por: Cranney, Caleb, et al.
Publicado: (2025)
Scaling Graph Transformers: A Comparative Study of Sparse and Dense Attention
por: Dimitrov, Leon
Publicado: (2025)
por: Dimitrov, Leon
Publicado: (2025)
Horizon-wise Learning Paradigm Promotes Gene Splicing Identification
por: Li, Qi-Jie, et al.
Publicado: (2024)
por: Li, Qi-Jie, et al.
Publicado: (2024)
CITRAS: Covariate-Informed Transformer for Time Series Forecasting
por: Yamaguchi, Yosuke, et al.
Publicado: (2025)
por: Yamaguchi, Yosuke, et al.
Publicado: (2025)
NoiseFormer -- Noise Diffused Symmetric Attention Transformer
por: Kumar, Phani, et al.
Publicado: (2026)
por: Kumar, Phani, et al.
Publicado: (2026)
Ejemplares similares
-
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
por: Adhikari, Rabin
Publicado: (2025) -
Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling
por: Zhao, Weijie, et al.
Publicado: (2026) -
WLFM: A Well-Logs Foundation Model for Multi-Task and Cross-Well Geological Interpretation
por: Qi, Zhenyu, et al.
Publicado: (2025) -
Neurocircuitry-Inspired Hierarchical Graph Causal Attention Networks for Explainable Depression Identification
por: Chen, Weidao, et al.
Publicado: (2025) -
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
por: Bu, Rui, et al.
Publicado: (2025)