Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yukun, Zhou, Xueqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Where to Add PDE Diffusion in Transformers
by: Zhang, Yukun, et al.
Published: (2025)
by: Zhang, Yukun, et al.
Published: (2025)
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
by: Zhang, Yukun, et al.
Published: (2024)
by: Zhang, Yukun, et al.
Published: (2024)
LMILAtt: A Deep Learning Model for Depression Detection from Social Media Users Enhanced by Multi-Instance Learning Based on Attention Mechanism
by: Yang, Yukun
Published: (2025)
by: Yang, Yukun
Published: (2025)
Physics-Guided Transformer (PGT): Physics-Aware Attention Mechanism for PINNs
by: Zeraatkar, Ehsan, et al.
Published: (2026)
by: Zeraatkar, Ehsan, et al.
Published: (2026)
Knowledge-enhanced Transformer for Multivariate Long Sequence Time-series Forecasting
by: Kakde, Shubham Tanaji, et al.
Published: (2024)
by: Kakde, Shubham Tanaji, et al.
Published: (2024)
Fourier-Mixed Window Attention: Accelerating Informer for Long Sequence Time-Series Forecasting
by: Tran, Nhat Thanh, et al.
Published: (2023)
by: Tran, Nhat Thanh, et al.
Published: (2023)
Enhancing Transformer-based models for Long Sequence Time Series Forecasting via Structured Matrix
by: Zhang, Zhicheng, et al.
Published: (2024)
by: Zhang, Zhicheng, et al.
Published: (2024)
Unisolver: PDE-Conditional Transformers Towards Universal Neural PDE Solvers
by: Zhou, Hang, et al.
Published: (2024)
by: Zhou, Hang, et al.
Published: (2024)
Long Input Sequence Network for Long Time Series Forecasting
by: Ma, Chao, et al.
Published: (2024)
by: Ma, Chao, et al.
Published: (2024)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
by: Dentamaro, Vincenzo
Published: (2025)
by: Dentamaro, Vincenzo
Published: (2025)
Revealing the Attention Floating Mechanism in Masked Diffusion Models
by: Dai, Xin, et al.
Published: (2026)
by: Dai, Xin, et al.
Published: (2026)
PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
by: Sun, Tian, et al.
Published: (2025)
by: Sun, Tian, et al.
Published: (2025)
AttnGen: Attention-Guided Saliency Learning for Interpretable Genomic Sequence Classification
by: Nia, Rayhaneh Shabani, et al.
Published: (2026)
by: Nia, Rayhaneh Shabani, et al.
Published: (2026)
XicorAttention: Time Series Transformer Using Attention with Nonlinear Correlation
by: Kimura, Daichi, et al.
Published: (2025)
by: Kimura, Daichi, et al.
Published: (2025)
PeSANet: Physics-encoded Spectral Attention Network for Simulating PDE-Governed Complex Systems
by: Wan, Han, et al.
Published: (2025)
by: Wan, Han, et al.
Published: (2025)
Unveiling LLM Mechanisms Through Neural ODEs and Control Theory
by: Zhang, Yukun, et al.
Published: (2024)
by: Zhang, Yukun, et al.
Published: (2024)
Star Attention: Efficient LLM Inference over Long Sequences
by: Acharya, Shantanu, et al.
Published: (2024)
by: Acharya, Shantanu, et al.
Published: (2024)
Attention in Constant Time: Vashista Sparse Attention for Long-Context Decoding with Exponential Guarantees
by: Nobaub, Vashista
Published: (2026)
by: Nobaub, Vashista
Published: (2026)
CATO: Charted Attention for Neural PDE Operators
by: Cheng, Chun-Wun, et al.
Published: (2026)
by: Cheng, Chun-Wun, et al.
Published: (2026)
PSformer: Parameter-efficient Transformer with Segment Attention for Time Series Forecasting
by: Wang, Yanlong, et al.
Published: (2024)
by: Wang, Yanlong, et al.
Published: (2024)
GSA-Forecaster: Forecasting Graph-Based Time-Dependent Data with Graph Sequence Attention
by: Li, Yang, et al.
Published: (2021)
by: Li, Yang, et al.
Published: (2021)
Space-Time Continuous PDE Forecasting using Equivariant Neural Fields
by: Knigge, David M., et al.
Published: (2024)
by: Knigge, David M., et al.
Published: (2024)
CFO: Learning Continuous-Time PDE Dynamics via Flow-Matched Neural Operators
by: Hou, Xianglong, et al.
Published: (2025)
by: Hou, Xianglong, et al.
Published: (2025)
InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model
by: Wang, Youjin, et al.
Published: (2026)
by: Wang, Youjin, et al.
Published: (2026)
Towards Generalizable PDE Dynamics Forecasting via Physics-Guided Invariant Learning
by: Li, Siyang, et al.
Published: (2025)
by: Li, Siyang, et al.
Published: (2025)
Introducing Spectral Attention for Long-Range Dependency in Time Series Forecasting
by: Kang, Bong Gyun, et al.
Published: (2024)
by: Kang, Bong Gyun, et al.
Published: (2024)
DiffusionRollout: Uncertainty-Aware Rollout Planning in Long-Horizon PDE Solving
by: Yoo, Seungwoo, et al.
Published: (2026)
by: Yoo, Seungwoo, et al.
Published: (2026)
Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive Forecasting
by: Lu, Jiecheng, et al.
Published: (2025)
by: Lu, Jiecheng, et al.
Published: (2025)
ContiFormer: Continuous-Time Transformer for Irregular Time Series Modeling
by: Chen, Yuqi, et al.
Published: (2024)
by: Chen, Yuqi, et al.
Published: (2024)
Transforming Multidimensional Time Series into Interpretable Event Sequences for Advanced Data Mining
by: Yan, Xu, et al.
Published: (2024)
by: Yan, Xu, et al.
Published: (2024)
Rough Transformers for Continuous and Efficient Time-Series Modelling
by: Moreno-Pino, Fernando, et al.
Published: (2024)
by: Moreno-Pino, Fernando, et al.
Published: (2024)
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
PeriodNet: Boosting the Potential of Attention Mechanism for Time Series Forecasting
by: Zhao, Bowen, et al.
Published: (2025)
by: Zhao, Bowen, et al.
Published: (2025)
DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
Triple Attention Transformer Architecture for Time-Dependent Concrete Creep Prediction
by: Dokduea, Warayut, et al.
Published: (2025)
by: Dokduea, Warayut, et al.
Published: (2025)
Linear Attention for Efficient Bidirectional Sequence Modeling
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
PDE-regularized Dynamics-informed Diffusion with Uncertainty-aware Filtering for Long-Horizon Dynamics
by: Baeg, Min Young, et al.
Published: (2026)
by: Baeg, Min Young, et al.
Published: (2026)
DGTN: Graph-Enhanced Transformer with Diffusive Attention Gating Mechanism for Enzyme DDG Prediction
by: Lin, Abigail
Published: (2025)
by: Lin, Abigail
Published: (2025)
When Attention Beats Fourier: Multi-Scale Transformers for PDE Solving on Irregular Domains
by: Yee, Brandon, et al.
Published: (2026)
by: Yee, Brandon, et al.
Published: (2026)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
by: Zhou, Ruijie, et al.
Published: (2026)
by: Zhou, Ruijie, et al.
Published: (2026)
Similar Items
-
Where to Add PDE Diffusion in Transformers
by: Zhang, Yukun, et al.
Published: (2025) -
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
by: Zhang, Yukun, et al.
Published: (2024) -
LMILAtt: A Deep Learning Model for Depression Detection from Social Media Users Enhanced by Multi-Instance Learning Based on Attention Mechanism
by: Yang, Yukun
Published: (2025) -
Physics-Guided Transformer (PGT): Physics-Aware Attention Mechanism for PINNs
by: Zeraatkar, Ehsan, et al.
Published: (2026) -
Knowledge-enhanced Transformer for Multivariate Long Sequence Time-series Forecasting
by: Kakde, Shubham Tanaji, et al.
Published: (2024)