Quantitative Bounds for Length Generalization in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Izzo, Zachary, Nichani, Eshaan, Lee, Jason D. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Transformers Learn Causal Structure with Gradient Descent
by: Nichani, Eshaan, et al.
Published: (2024)
by: Nichani, Eshaan, et al.
Published: (2024)
Understanding Factual Recall in Transformers via Associative Memories
by: Nichani, Eshaan, et al.
Published: (2024)
by: Nichani, Eshaan, et al.
Published: (2024)
Fine-Tuning Dynamics of In-Context Factual Recall in Transformers
by: Huang, Ruomin, et al.
Published: (2026)
by: Huang, Ruomin, et al.
Published: (2026)
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
by: Nichani, Eshaan, et al.
Published: (2023)
by: Nichani, Eshaan, et al.
Published: (2023)
Emergence and scaling laws in SGD learning of shallow neural networks
by: Ren, Yunwei, et al.
Published: (2025)
by: Ren, Yunwei, et al.
Published: (2025)
On the Statistical Query Complexity of Learning Semiautomata: a Random Walk Approach
by: Giapitzakis, George, et al.
Published: (2025)
by: Giapitzakis, George, et al.
Published: (2025)
Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks
by: Fu, Hengyu, et al.
Published: (2024)
by: Fu, Hengyu, et al.
Published: (2024)
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
by: Kim, Juno, et al.
Published: (2026)
by: Kim, Juno, et al.
Published: (2026)
Learning Compositional Functions with Transformers from Easy-to-Hard Data
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval
by: Barnfield, Nicholas, et al.
Published: (2026)
by: Barnfield, Nicholas, et al.
Published: (2026)
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
Length Generalization Bounds for Transformers
by: Yang, Andy, et al.
Published: (2026)
by: Yang, Andy, et al.
Published: (2026)
Looped Transformers for Length Generalization
by: Fan, Ying, et al.
Published: (2024)
by: Fan, Ying, et al.
Published: (2024)
Subgroup Discovery with the Cox Model
by: Izzo, Zachary, et al.
Published: (2025)
by: Izzo, Zachary, et al.
Published: (2025)
Queue Length Regret Bounds for Contextual Queueing Bandits
by: Bae, Seoungbin, et al.
Published: (2026)
by: Bae, Seoungbin, et al.
Published: (2026)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
by: Lee, Nayoung, et al.
Published: (2025)
by: Lee, Nayoung, et al.
Published: (2025)
On Vanishing Variance in Transformer Length Generalization
by: Li, Ruining, et al.
Published: (2025)
by: Li, Ruining, et al.
Published: (2025)
A Formal Framework for Understanding Length Generalization in Transformers
by: Huang, Xinting, et al.
Published: (2024)
by: Huang, Xinting, et al.
Published: (2024)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
Arbitrary-Length Generalization for Addition in a Tiny Transformer
by: Patriota, Alexandre Galvao
Published: (2024)
by: Patriota, Alexandre Galvao
Published: (2024)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
by: Phan, Buu, et al.
Published: (2025)
by: Phan, Buu, et al.
Published: (2025)
Position Encoding with Random Float Sampling Enhances Length Generalization of Transformers
by: Shimizu, Atsushi, et al.
Published: (2026)
by: Shimizu, Atsushi, et al.
Published: (2026)
Sharper Generalization Bounds for Transformer
by: Li, Yawen, et al.
Published: (2026)
by: Li, Yawen, et al.
Published: (2026)
Transformers Can Achieve Length Generalization But Not Robustly
by: Zhou, Yongchao, et al.
Published: (2024)
by: Zhou, Yongchao, et al.
Published: (2024)
From Interpolation to Extrapolation: Complete Length Generalization for Arithmetic Transformers
by: Duan, Shaoxiong, et al.
Published: (2023)
by: Duan, Shaoxiong, et al.
Published: (2023)
Spectrum-Adaptive Generalization Bounds for Trained Deep Transformers
by: Sakai, Mana, et al.
Published: (2026)
by: Sakai, Mana, et al.
Published: (2026)
Guidance and Control Networks with Periodic Activation Functions
by: Origer, Sebastien, et al.
Published: (2024)
by: Origer, Sebastien, et al.
Published: (2024)
Generalization Bounds for Transformer Channel Decoders
by: Zhang, Qinshan, et al.
Published: (2026)
by: Zhang, Qinshan, et al.
Published: (2026)
Learning and Transferring Sparse Contextual Bigrams with Linear Transformers
by: Ren, Yunwei, et al.
Published: (2024)
by: Ren, Yunwei, et al.
Published: (2024)
Mesh-Informed Neural Operator : A Transformer Generative Approach
by: Shi, Yaozhong, et al.
Published: (2025)
by: Shi, Yaozhong, et al.
Published: (2025)
Non-Asymptotic Length Generalization
by: Chen, Thomas, et al.
Published: (2025)
by: Chen, Thomas, et al.
Published: (2025)
Does Privacy Always Harm Fairness? Data-Dependent Trade-offs via Chernoff Information Neural Estimation
by: Nichani, Arjun, et al.
Published: (2026)
by: Nichani, Arjun, et al.
Published: (2026)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
by: Huang, Jianhao, et al.
Published: (2025)
by: Huang, Jianhao, et al.
Published: (2025)
Group Relative Augmentation for Data Efficient Action Detection
by: Patel, Deep Anil, et al.
Published: (2025)
by: Patel, Deep Anil, et al.
Published: (2025)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
The Generative Leap: Sharp Sample Complexity for Efficiently Learning Gaussian Multi-Index Models
by: Damian, Alex, et al.
Published: (2025)
by: Damian, Alex, et al.
Published: (2025)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
A Tight Lower Bound for Non-stochastic Multi-armed Bandits with Expert Advice
by: Chase, Zachary, et al.
Published: (2025)
by: Chase, Zachary, et al.
Published: (2025)
Optimal Mistake Bounds for Transductive Online Learning
by: Chase, Zachary, et al.
Published: (2025)
by: Chase, Zachary, et al.
Published: (2025)
Similar Items
-
How Transformers Learn Causal Structure with Gradient Descent
by: Nichani, Eshaan, et al.
Published: (2024) -
Understanding Factual Recall in Transformers via Associative Memories
by: Nichani, Eshaan, et al.
Published: (2024) -
Fine-Tuning Dynamics of In-Context Factual Recall in Transformers
by: Huang, Ruomin, et al.
Published: (2026) -
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
by: Nichani, Eshaan, et al.
Published: (2023) -
Emergence and scaling laws in SGD learning of shallow neural networks
by: Ren, Yunwei, et al.
Published: (2025)