Short-Range Dependency Effects on Transformer Instability and a Decomposed Attention Solution
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Hajra, Suvadeep |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Decomposable Transformer Point Processes
von: Panos, Aristeidis
Veröffentlicht: (2024)
von: Panos, Aristeidis
Veröffentlicht: (2024)
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
von: Hajra, Suvadeep, et al.
Veröffentlicht: (2026)
von: Hajra, Suvadeep, et al.
Veröffentlicht: (2026)
Introducing Spectral Attention for Long-Range Dependency in Time Series Forecasting
von: Kang, Bong Gyun, et al.
Veröffentlicht: (2024)
von: Kang, Bong Gyun, et al.
Veröffentlicht: (2024)
Short-Range Oversquashing
von: Mishayev, Yaaqov, et al.
Veröffentlicht: (2025)
von: Mishayev, Yaaqov, et al.
Veröffentlicht: (2025)
Decomposing Attention To Find Context-Sensitive Neurons
von: Gibson, Alex
Veröffentlicht: (2025)
von: Gibson, Alex
Veröffentlicht: (2025)
Hybrid Focal and Full-Range Attention Based Graph Transformers
von: Zhu, Minhong, et al.
Veröffentlicht: (2023)
von: Zhu, Minhong, et al.
Veröffentlicht: (2023)
Transparency in Sleep Staging: Deep Learning Method for EEG Sleep Stage Classification with Model Interpretability
von: Sharma, Shivam, et al.
Veröffentlicht: (2023)
von: Sharma, Shivam, et al.
Veröffentlicht: (2023)
Decomposing Global Feature Effects Based on Feature Interactions
von: Herbinger, Julia, et al.
Veröffentlicht: (2023)
von: Herbinger, Julia, et al.
Veröffentlicht: (2023)
The Effect of Attention Head Count on Transformer Approximation
von: Yu, Penghao, et al.
Veröffentlicht: (2025)
von: Yu, Penghao, et al.
Veröffentlicht: (2025)
Triple Attention Transformer Architecture for Time-Dependent Concrete Creep Prediction
von: Dokduea, Warayut, et al.
Veröffentlicht: (2025)
von: Dokduea, Warayut, et al.
Veröffentlicht: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
von: Lou, Chao, et al.
Veröffentlicht: (2024)
von: Lou, Chao, et al.
Veröffentlicht: (2024)
DOLCE: Decomposing Off-Policy Evaluation/Learning into Lagged and Current Effects
von: Tamano, Shu
Veröffentlicht: (2025)
von: Tamano, Shu
Veröffentlicht: (2025)
Integration of Mamba and Transformer -- MAT for Long-Short Range Time Series Forecasting with Application to Weather Dynamics
von: Zhang, Wenqing, et al.
Veröffentlicht: (2024)
von: Zhang, Wenqing, et al.
Veröffentlicht: (2024)
Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2025)
Extracting Cause-Effect Pairs from a Sentence with a Dependency-Aware Transformer Model
von: Kabir, Md Ahsanul, et al.
Veröffentlicht: (2025)
von: Kabir, Md Ahsanul, et al.
Veröffentlicht: (2025)
Learning Long-Range Dependencies with Temporal Predictive Coding
von: Potter, Tom, et al.
Veröffentlicht: (2026)
von: Potter, Tom, et al.
Veröffentlicht: (2026)
Decomposable Neuro Symbolic Regression
von: Morales, Giorgio, et al.
Veröffentlicht: (2025)
von: Morales, Giorgio, et al.
Veröffentlicht: (2025)
DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
Learning Long Range Dependencies on Graphs via Random Walks
von: Chen, Dexiong, et al.
Veröffentlicht: (2024)
von: Chen, Dexiong, et al.
Veröffentlicht: (2024)
AI Generalisation Gap In Comorbid Sleep Disorder Staging
von: Bose, Saswata, et al.
Veröffentlicht: (2026)
von: Bose, Saswata, et al.
Veröffentlicht: (2026)
On the Effect of Instability on Learning Continuous-Time Linear Control Systems
von: Hafshejani, Reza Sadeghi, et al.
Veröffentlicht: (2024)
von: Hafshejani, Reza Sadeghi, et al.
Veröffentlicht: (2024)
Decomposing Prediction Mechanisms for In-Context Recall
von: Daniels, Sultan, et al.
Veröffentlicht: (2025)
von: Daniels, Sultan, et al.
Veröffentlicht: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
von: Engels, Joshua, et al.
Veröffentlicht: (2024)
von: Engels, Joshua, et al.
Veröffentlicht: (2024)
Decomposing the Depth Profile of Fine-Tuning
von: Billa, Jayadev
Veröffentlicht: (2026)
von: Billa, Jayadev
Veröffentlicht: (2026)
Decomposing Gaussians with Unknown Covariance
von: Dharamshi, Ameer, et al.
Veröffentlicht: (2024)
von: Dharamshi, Ameer, et al.
Veröffentlicht: (2024)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
ParallelTime: Dynamically Weighting the Balance of Short- and Long-Term Temporal Dependencies
von: Katav, Itay, et al.
Veröffentlicht: (2025)
von: Katav, Itay, et al.
Veröffentlicht: (2025)
Leveraging Discrete Function Decomposability for Scientific Design
von: Bowden, James C., et al.
Veröffentlicht: (2025)
von: Bowden, James C., et al.
Veröffentlicht: (2025)
Decomposing Task Vectors for Refined Model Editing
von: Damirchi, Hamed, et al.
Veröffentlicht: (2025)
von: Damirchi, Hamed, et al.
Veröffentlicht: (2025)
DEMAU: Decompose, Explore, Model and Analyse Uncertainties
von: Hoarau, Arthur, et al.
Veröffentlicht: (2024)
von: Hoarau, Arthur, et al.
Veröffentlicht: (2024)
On the Mirage of Long-Range Dependency, with an Application to Integer Multiplication
von: Wei, Zichao
Veröffentlicht: (2026)
von: Wei, Zichao
Veröffentlicht: (2026)
Batch Normalization Decomposed
von: Nachum, Ido, et al.
Veröffentlicht: (2024)
von: Nachum, Ido, et al.
Veröffentlicht: (2024)
Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
von: Kim, Bum Jun, et al.
Veröffentlicht: (2026)
von: Kim, Bum Jun, et al.
Veröffentlicht: (2026)
Transformer Reconstructed with Dynamic Value Attention
von: Wang, Xiaowei
Veröffentlicht: (2025)
von: Wang, Xiaowei
Veröffentlicht: (2025)
Cottention: Linear Transformers With Cosine Attention
von: Mongaras, Gabriel, et al.
Veröffentlicht: (2024)
von: Mongaras, Gabriel, et al.
Veröffentlicht: (2024)
Transformers with Sparse Attention for Granger Causality
von: Mahesh, Riya, et al.
Veröffentlicht: (2024)
von: Mahesh, Riya, et al.
Veröffentlicht: (2024)
Preconditioned Attention: Enhancing Efficiency in Transformers
von: Saratchandran, Hemanth
Veröffentlicht: (2026)
von: Saratchandran, Hemanth
Veröffentlicht: (2026)
Graph External Attention Enhanced Transformer
von: Liang, Jianqing, et al.
Veröffentlicht: (2024)
von: Liang, Jianqing, et al.
Veröffentlicht: (2024)
Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache
von: Dehghankar, Mohsen, et al.
Veröffentlicht: (2026)
von: Dehghankar, Mohsen, et al.
Veröffentlicht: (2026)
Decomposing multimodal embedding spaces with group-sparse autoencoders
von: Kaushik, Chiraag, et al.
Veröffentlicht: (2026)
von: Kaushik, Chiraag, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Decomposable Transformer Point Processes
von: Panos, Aristeidis
Veröffentlicht: (2024) -
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
von: Hajra, Suvadeep, et al.
Veröffentlicht: (2026) -
Introducing Spectral Attention for Long-Range Dependency in Time Series Forecasting
von: Kang, Bong Gyun, et al.
Veröffentlicht: (2024) -
Short-Range Oversquashing
von: Mishayev, Yaaqov, et al.
Veröffentlicht: (2025) -
Decomposing Attention To Find Context-Sensitive Neurons
von: Gibson, Alex
Veröffentlicht: (2025)