Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yukun, Zhou, Xueqing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908735891308544
author Zhang, Yukun
Zhou, Xueqing
author_facet Zhang, Yukun
Zhou, Xueqing
contents We propose a novel framework, Continuous_Time Attention, which infuses partial differential equations (PDEs) into the Transformer's attention mechanism to address the challenges of extremely long input sequences. Instead of relying solely on a static attention matrix, we allow attention weights to evolve over a pseudo_time dimension via diffusion, wave, or reaction_diffusion dynamics. This mechanism systematically smooths local noise, enhances long_range dependencies, and stabilizes gradient flow. Theoretically, our analysis shows that PDE_based attention leads to better optimization landscapes and polynomial rather than exponential decay of distant interactions. Empirically, we benchmark our method on diverse experiments_demonstrating consistent gains over both standard and specialized long sequence Transformer variants. Our findings highlight the potential of PDE_based formulations to enrich attention mechanisms with continuous_time dynamics and global coherence.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20666
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
Zhang, Yukun
Zhou, Xueqing
Machine Learning
Artificial Intelligence
We propose a novel framework, Continuous_Time Attention, which infuses partial differential equations (PDEs) into the Transformer's attention mechanism to address the challenges of extremely long input sequences. Instead of relying solely on a static attention matrix, we allow attention weights to evolve over a pseudo_time dimension via diffusion, wave, or reaction_diffusion dynamics. This mechanism systematically smooths local noise, enhances long_range dependencies, and stabilizes gradient flow. Theoretically, our analysis shows that PDE_based attention leads to better optimization landscapes and polynomial rather than exponential decay of distant interactions. Empirically, we benchmark our method on diverse experiments_demonstrating consistent gains over both standard and specialized long sequence Transformer variants. Our findings highlight the potential of PDE_based formulations to enrich attention mechanisms with continuous_time dynamics and global coherence.
title Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.20666