A Formal Framework for Understanding Length Generalization in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Xinting, Yang, Andy, Bhattamishra, Satwik, Sarrof, Yash, Krebs, Andreas, Zhou, Hattie, Nakkiran, Preetum, Hahn, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
by: Huang, Xinting, et al.
Published: (2026)
by: Huang, Xinting, et al.
Published: (2026)
Step-by-Step Diffusion: An Elementary Tutorial
by: Nakkiran, Preetum, et al.
Published: (2024)
by: Nakkiran, Preetum, et al.
Published: (2024)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
by: Bhattamishra, Satwik, et al.
Published: (2024)
by: Bhattamishra, Satwik, et al.
Published: (2024)
The Expressive Capacity of State Space Models: A Formal Language Perspective
by: Sarrof, Yash, et al.
Published: (2024)
by: Sarrof, Yash, et al.
Published: (2024)
Provably Learning Attention with Queries
by: Bhattamishra, Satwik, et al.
Published: (2026)
by: Bhattamishra, Satwik, et al.
Published: (2026)
Classifier-Free Guidance is a Predictor-Corrector
by: Bradley, Arwen, et al.
Published: (2024)
by: Bradley, Arwen, et al.
Published: (2024)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
by: Kraus, Oliver, et al.
Published: (2026)
by: Kraus, Oliver, et al.
Published: (2026)
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
by: Bhattamishra, Satwik, et al.
Published: (2025)
by: Bhattamishra, Satwik, et al.
Published: (2025)
Vanishing Gradients in Reinforcement Finetuning of Language Models
by: Razin, Noam, et al.
Published: (2023)
by: Razin, Noam, et al.
Published: (2023)
When is Multicalibration Post-Processing Necessary?
by: Hansen, Dutch, et al.
Published: (2024)
by: Hansen, Dutch, et al.
Published: (2024)
On the Ability of Transformers to Verify Plans
by: Sarrof, Yash, et al.
Published: (2026)
by: Sarrof, Yash, et al.
Published: (2026)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
by: Jobanputra, Mayank, et al.
Published: (2025)
by: Jobanputra, Mayank, et al.
Published: (2025)
Benefits and Limitations of Communication in Multi-Agent Reasoning
by: Rizvi-Martel, Michael, et al.
Published: (2025)
by: Rizvi-Martel, Michael, et al.
Published: (2025)
The Transformer Cookbook
by: Yang, Andy, et al.
Published: (2025)
by: Yang, Andy, et al.
Published: (2025)
Mechanisms of Projective Composition of Diffusion Models
by: Bradley, Arwen, et al.
Published: (2025)
by: Bradley, Arwen, et al.
Published: (2025)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
by: Nakkiran, Preetum, et al.
Published: (2025)
by: Nakkiran, Preetum, et al.
Published: (2025)
Composition and Control with Distilled Energy Diffusion Models and Sequential Monte Carlo
by: Thornton, James, et al.
Published: (2025)
by: Thornton, James, et al.
Published: (2025)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
by: Huang, Xinting, et al.
Published: (2025)
by: Huang, Xinting, et al.
Published: (2025)
How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
by: Littwin, Etai, et al.
Published: (2024)
by: Littwin, Etai, et al.
Published: (2024)
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
by: Amiri, Alireza, et al.
Published: (2025)
by: Amiri, Alireza, et al.
Published: (2025)
InversionView: A General-Purpose Method for Reading Information from Neural Activations
by: Huang, Xinting, et al.
Published: (2024)
by: Huang, Xinting, et al.
Published: (2024)
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
by: Mallinar, Neil, et al.
Published: (2022)
by: Mallinar, Neil, et al.
Published: (2022)
Inshrinkerator: Compressing Deep Learning Training Checkpoints via Dynamic Quantization
by: Agrawal, Amey, et al.
Published: (2023)
by: Agrawal, Amey, et al.
Published: (2023)
Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B
by: Bakalova, Aleksandra, et al.
Published: (2025)
by: Bakalova, Aleksandra, et al.
Published: (2025)
Length Generalization Bounds for Transformers
by: Yang, Andy, et al.
Published: (2026)
by: Yang, Andy, et al.
Published: (2026)
Normalizing Flows are Capable Generative Models
by: Zhai, Shuangfei, et al.
Published: (2024)
by: Zhai, Shuangfei, et al.
Published: (2024)
On Vanishing Variance in Transformer Length Generalization
by: Li, Ruining, et al.
Published: (2025)
by: Li, Ruining, et al.
Published: (2025)
Looped Transformers for Length Generalization
by: Fan, Ying, et al.
Published: (2024)
by: Fan, Ying, et al.
Published: (2024)
MALT Diffusion: Memory-Augmented Latent Transformers for Any-Length Video Generation
by: Yu, Sihyun, et al.
Published: (2025)
by: Yu, Sihyun, et al.
Published: (2025)
Quantitative Bounds for Length Generalization in Transformers
by: Izzo, Zachary, et al.
Published: (2025)
by: Izzo, Zachary, et al.
Published: (2025)
Gating Enables Curvature: A Geometric Expressivity Gap in Attention
by: Bathula, Satwik, et al.
Published: (2026)
by: Bathula, Satwik, et al.
Published: (2026)
Understanding and Improving Length Generalization in Recurrent Models
by: Ruiz, Ricardo Buitrago, et al.
Published: (2025)
by: Ruiz, Ricardo Buitrago, et al.
Published: (2025)
Why are Sensitive Functions Hard for Transformers?
by: Hahn, Michael, et al.
Published: (2024)
by: Hahn, Michael, et al.
Published: (2024)
Trace Length is a Simple Uncertainty Signal in Reasoning Models
by: Devic, Siddartha, et al.
Published: (2025)
by: Devic, Siddartha, et al.
Published: (2025)
Transformers Can Achieve Length Generalization But Not Robustly
by: Zhou, Yongchao, et al.
Published: (2024)
by: Zhou, Yongchao, et al.
Published: (2024)
STIQ: Safeguarding Training and Inferencing of Quantum Neural Networks from Untrusted Cloud
by: Kundu, Satwik, et al.
Published: (2024)
by: Kundu, Satwik, et al.
Published: (2024)
Inverse-Transpilation: Reverse-Engineering Quantum Compiler Optimization Passes from Circuit Snapshots
by: Kundu, Satwik, et al.
Published: (2025)
by: Kundu, Satwik, et al.
Published: (2025)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
Arbitrary-Length Generalization for Addition in a Tiny Transformer
by: Patriota, Alexandre Galvao
Published: (2024)
by: Patriota, Alexandre Galvao
Published: (2024)
Security Concerns in Quantum Machine Learning as a Service
by: Kundu, Satwik, et al.
Published: (2024)
by: Kundu, Satwik, et al.
Published: (2024)
Similar Items
-
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
by: Huang, Xinting, et al.
Published: (2026) -
Step-by-Step Diffusion: An Elementary Tutorial
by: Nakkiran, Preetum, et al.
Published: (2024) -
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
by: Bhattamishra, Satwik, et al.
Published: (2024) -
The Expressive Capacity of State Space Models: A Formal Language Perspective
by: Sarrof, Yash, et al.
Published: (2024) -
Provably Learning Attention with Queries
by: Bhattamishra, Satwik, et al.
Published: (2026)