Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Phan, Buu, Ebrahimi, Reza, Haresh, Sanjay, Memisevic, Roland |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ClevrSkills: Compositional Language and Visual Reasoning in Robotics
von: Haresh, Sanjay, et al.
Veröffentlicht: (2024)
von: Haresh, Sanjay, et al.
Veröffentlicht: (2024)
Revisiting Bi-Linear State Transitions in Recurrent Neural Networks
von: Ebrahimi, M. Reza, et al.
Veröffentlicht: (2025)
von: Ebrahimi, M. Reza, et al.
Veröffentlicht: (2025)
On the "Induction Bias" in Sequence Models
von: Ebrahimi, M. Reza, et al.
Veröffentlicht: (2026)
von: Ebrahimi, M. Reza, et al.
Veröffentlicht: (2026)
Channel Simulation and Distributed Compression with Ensemble Rejection Sampling
von: Phan, Buu, et al.
Veröffentlicht: (2025)
von: Phan, Buu, et al.
Veröffentlicht: (2025)
List-Level Distribution Coupling with Applications to Speculative Decoding and Lossy Compression
von: Rowan, Joseph, et al.
Veröffentlicht: (2025)
von: Rowan, Joseph, et al.
Veröffentlicht: (2025)
Attention as an RNN
von: Feng, Leo, et al.
Veröffentlicht: (2024)
von: Feng, Leo, et al.
Veröffentlicht: (2024)
Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks
von: Haresh, Sanjay, et al.
Veröffentlicht: (2026)
von: Haresh, Sanjay, et al.
Veröffentlicht: (2026)
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
von: Phan, Buu, et al.
Veröffentlicht: (2025)
von: Phan, Buu, et al.
Veröffentlicht: (2025)
Multi-Draft Speculative Sampling: Canonical Decomposition and Theoretical Limits
von: Khisti, Ashish, et al.
Veröffentlicht: (2024)
von: Khisti, Ashish, et al.
Veröffentlicht: (2024)
Your Context Is Not an Array: Unveiling Random Access Limitations in Transformers
von: Ebrahimi, MohammadReza, et al.
Veröffentlicht: (2024)
von: Ebrahimi, MohammadReza, et al.
Veröffentlicht: (2024)
Replacing thinking with tool usage enables reasoning in small language models
von: Rainone, Corrado, et al.
Veröffentlicht: (2025)
von: Rainone, Corrado, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Tokenization Bias in Language Models
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
Aligning Robot Navigation Behaviors with Human Intentions and Preferences
von: Karnan, Haresh
Veröffentlicht: (2024)
von: Karnan, Haresh
Veröffentlicht: (2024)
Look, Remember and Reason: Grounded reasoning in videos with language models
von: Bhattacharyya, Apratim, et al.
Veröffentlicht: (2023)
von: Bhattacharyya, Apratim, et al.
Veröffentlicht: (2023)
Enhancing Time Series Forecasting with Fuzzy Attention-Integrated Transformers
von: Chakraborty, Sanjay, et al.
Veröffentlicht: (2025)
von: Chakraborty, Sanjay, et al.
Veröffentlicht: (2025)
Attention Augmented GNN RNN-Attention Models for Advanced Cybersecurity Intrusion Detection
von: Biradar, Jayant, et al.
Veröffentlicht: (2025)
von: Biradar, Jayant, et al.
Veröffentlicht: (2025)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
von: Leng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Leng, Jiaqi, et al.
Veröffentlicht: (2025)
TRA: Better Length Generalisation with Threshold Relative Attention
von: Opper, Mattia, et al.
Veröffentlicht: (2025)
von: Opper, Mattia, et al.
Veröffentlicht: (2025)
Looped Transformers for Length Generalization
von: Fan, Ying, et al.
Veröffentlicht: (2024)
von: Fan, Ying, et al.
Veröffentlicht: (2024)
Learning to Forget with Information Divergence Reweighted Objectives for Noisy Labels
von: Birrell, Jeremiah, et al.
Veröffentlicht: (2025)
von: Birrell, Jeremiah, et al.
Veröffentlicht: (2025)
Optimal Transport Regularized Divergences: Application to Adversarial Robustness
von: Birrell, Jeremiah, et al.
Veröffentlicht: (2023)
von: Birrell, Jeremiah, et al.
Veröffentlicht: (2023)
Integrating Quantum-Classical Attention in Patch Transformers for Enhanced Time Series Forecasting
von: Chakraborty, Sanjay, et al.
Veröffentlicht: (2025)
von: Chakraborty, Sanjay, et al.
Veröffentlicht: (2025)
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)
Quantitative Bounds for Length Generalization in Transformers
von: Izzo, Zachary, et al.
Veröffentlicht: (2025)
von: Izzo, Zachary, et al.
Veröffentlicht: (2025)
When Learning Hurts: Fixed-Pole RNN for Real-Time Online Training
von: Morgan, Alexander, et al.
Veröffentlicht: (2026)
von: Morgan, Alexander, et al.
Veröffentlicht: (2026)
Understanding and Improving Length Generalization in Recurrent Models
von: Ruiz, Ricardo Buitrago, et al.
Veröffentlicht: (2025)
von: Ruiz, Ricardo Buitrago, et al.
Veröffentlicht: (2025)
The Norm-Separation Delay Law of Grokking: A First-Principles Theory of Delayed Generalization
von: Khanh, Truong Xuan, et al.
Veröffentlicht: (2026)
von: Khanh, Truong Xuan, et al.
Veröffentlicht: (2026)
Crash Severity Prediction Using Deep Learning Approaches: A Hybrid CNN-RNN Framework
von: Koohfar, Sahar
Veröffentlicht: (2025)
von: Koohfar, Sahar
Veröffentlicht: (2025)
On Vanishing Variance in Transformer Length Generalization
von: Li, Ruining, et al.
Veröffentlicht: (2025)
von: Li, Ruining, et al.
Veröffentlicht: (2025)
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
von: Beck, Maximilian, et al.
Veröffentlicht: (2025)
von: Beck, Maximilian, et al.
Veröffentlicht: (2025)
A Formal Framework for Understanding Length Generalization in Transformers
von: Huang, Xinting, et al.
Veröffentlicht: (2024)
von: Huang, Xinting, et al.
Veröffentlicht: (2024)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
von: Bendikas, Rokas, et al.
Veröffentlicht: (2025)
von: Bendikas, Rokas, et al.
Veröffentlicht: (2025)
Leakage and Second-Order Dynamics Improve Hippocampal RNN Replay
von: Casco-Rodriguez, Josue, et al.
Veröffentlicht: (2026)
von: Casco-Rodriguez, Josue, et al.
Veröffentlicht: (2026)
Spectral Conditioning of Attention Improves Transformer Performance
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2026)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2026)
HyMaTE: A Hybrid Mamba and Transformer Model for EHR Representation Learning
von: Mottalib, Md Mozaharul, et al.
Veröffentlicht: (2025)
von: Mottalib, Md Mozaharul, et al.
Veröffentlicht: (2025)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
On Self-Adaptive Perception Loss Function for Sequential Lossy Compression
von: Salehkalaibar, Sadaf, et al.
Veröffentlicht: (2025)
von: Salehkalaibar, Sadaf, et al.
Veröffentlicht: (2025)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
von: Cheng, Zicong, et al.
Veröffentlicht: (2026)
von: Cheng, Zicong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ClevrSkills: Compositional Language and Visual Reasoning in Robotics
von: Haresh, Sanjay, et al.
Veröffentlicht: (2024) -
Revisiting Bi-Linear State Transitions in Recurrent Neural Networks
von: Ebrahimi, M. Reza, et al.
Veröffentlicht: (2025) -
On the "Induction Bias" in Sequence Models
von: Ebrahimi, M. Reza, et al.
Veröffentlicht: (2026) -
Channel Simulation and Distributed Compression with Ensemble Rejection Sampling
von: Phan, Buu, et al.
Veröffentlicht: (2025) -
List-Level Distribution Coupling with Applications to Speculative Decoding and Lossy Compression
von: Rowan, Joseph, et al.
Veröffentlicht: (2025)