One-layer transformers fail to solve the induction heads task
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sanford, Clayton, Hsu, Daniel, Telgarsky, Matus |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transformers, parallel computation, and logarithmic depth
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
On Achieving Optimal Adversarial Test Error
von: Li, Justin D., et al.
Veröffentlicht: (2023)
von: Li, Justin D., et al.
Veröffentlicht: (2023)
Astral Space: Convex Analysis at Infinity
von: Dudík, Miroslav, et al.
Veröffentlicht: (2022)
von: Dudík, Miroslav, et al.
Veröffentlicht: (2022)
Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic Regression
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
Spectrum Extraction and Clipping for Implicitly Linear Layers
von: Boroojeny, Ali Ebrahimpour, et al.
Veröffentlicht: (2024)
von: Boroojeny, Ali Ebrahimpour, et al.
Veröffentlicht: (2024)
Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency
von: Wu, Jingfeng, et al.
Veröffentlicht: (2024)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2024)
Lower bounds for one-layer transformers that compute parity
von: Hsu, Daniel
Veröffentlicht: (2026)
von: Hsu, Daniel
Veröffentlicht: (2026)
Fast attention mechanisms: a tale of parallelism
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
Basic Inequalities for First-Order Optimization with Applications to Statistical Risk Analysis
von: Paik, Seunghoon, et al.
Veröffentlicht: (2025)
von: Paik, Seunghoon, et al.
Veröffentlicht: (2025)
Stream separation improves Bregman conditioning in transformers
von: Kerce, James Clayton
Veröffentlicht: (2026)
von: Kerce, James Clayton
Veröffentlicht: (2026)
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2025)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2025)
Attention layers provably solve single-location regression
von: Marion, Pierre, et al.
Veröffentlicht: (2024)
von: Marion, Pierre, et al.
Veröffentlicht: (2024)
Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers
von: Bechler-Speicher, Maya, et al.
Veröffentlicht: (2026)
von: Bechler-Speicher, Maya, et al.
Veröffentlicht: (2026)
Fine-tuned network relies on generic representation to solve unseen cognitive task
von: Lin, Dongyan
Veröffentlicht: (2024)
von: Lin, Dongyan
Veröffentlicht: (2024)
Small transformer architectures for task switching
von: Gros, Claudius
Veröffentlicht: (2025)
von: Gros, Claudius
Veröffentlicht: (2025)
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
Relational reasoning and inductive bias in transformers and large language models
von: Geerts, Jesse, et al.
Veröffentlicht: (2025)
von: Geerts, Jesse, et al.
Veröffentlicht: (2025)
A One-Inclusion Graph Approach to Multi-Group Learning
von: Bergam, Noah, et al.
Veröffentlicht: (2026)
von: Bergam, Noah, et al.
Veröffentlicht: (2026)
Next-Token Prediction and Regret Minimization
von: Mohri, Mehryar, et al.
Veröffentlicht: (2026)
von: Mohri, Mehryar, et al.
Veröffentlicht: (2026)
Best of Both Worlds: Advantages of Hybrid Graph Sequence Models
von: Behrouz, Ali, et al.
Veröffentlicht: (2024)
von: Behrouz, Ali, et al.
Veröffentlicht: (2024)
Provably learning a multi-head attention layer
von: Chen, Sitan, et al.
Veröffentlicht: (2024)
von: Chen, Sitan, et al.
Veröffentlicht: (2024)
Out-of-distribution generalization via composition: a lens through induction heads in Transformers
von: Song, Jiajun, et al.
Veröffentlicht: (2024)
von: Song, Jiajun, et al.
Veröffentlicht: (2024)
Tabular data generation with tensor contraction layers and transformers
von: Silva, Aníbal, et al.
Veröffentlicht: (2024)
von: Silva, Aníbal, et al.
Veröffentlicht: (2024)
Dynamic layer selection in decoder-only transformers
von: Glavas, Theodore, et al.
Veröffentlicht: (2024)
von: Glavas, Theodore, et al.
Veröffentlicht: (2024)
Understanding Transformer Reasoning Capabilities via Graph Algorithms
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
Depth-Width tradeoffs in Algorithmic Reasoning of Graph Tasks with Transformers
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
Trapped by simplicity: When Transformers fail to learn from noisy features
von: Peters, Evan, et al.
Veröffentlicht: (2026)
von: Peters, Evan, et al.
Veröffentlicht: (2026)
Adversarial Debiasing for Unbiased Parameter Recovery
von: Sanford, Luke C, et al.
Veröffentlicht: (2025)
von: Sanford, Luke C, et al.
Veröffentlicht: (2025)
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
From drift to adaptation to the failed ml model: Transfer Learning in Industrial MLOps
von: Ashraf, Waqar Muhammad, et al.
Veröffentlicht: (2026)
von: Ashraf, Waqar Muhammad, et al.
Veröffentlicht: (2026)
Propagation of Chaos in One-hidden-layer Neural Networks beyond Logarithmic Time
von: Glasgow, Margalit, et al.
Veröffentlicht: (2025)
von: Glasgow, Margalit, et al.
Veröffentlicht: (2025)
Unified CNNs and transformers underlying learning mechanism reveals multi-head attention modus vivendi
von: Koresh, Ella, et al.
Veröffentlicht: (2025)
von: Koresh, Ella, et al.
Veröffentlicht: (2025)
Shared-unique Features and Task-aware Prioritized Sampling on Multi-task Reinforcement Learning
von: Lin, Po-Shao, et al.
Veröffentlicht: (2024)
von: Lin, Po-Shao, et al.
Veröffentlicht: (2024)
Imitation of human motion achieves natural head movements for humanoid robots in an active-speaker detection task
von: Ding, Bosong, et al.
Veröffentlicht: (2024)
von: Ding, Bosong, et al.
Veröffentlicht: (2024)
Dimension lower bounds for linear approaches to function approximation
von: Hsu, Daniel
Veröffentlicht: (2025)
von: Hsu, Daniel
Veröffentlicht: (2025)
The Implicit Bias of Gradient Descent on Separable Multiclass Data
von: Ravi, Hrithik, et al.
Veröffentlicht: (2024)
von: Ravi, Hrithik, et al.
Veröffentlicht: (2024)
Improving One-class Recommendation with Multi-tasking on Various Preference Intensities
von: Shao, Chu-Jen, et al.
Veröffentlicht: (2024)
von: Shao, Chu-Jen, et al.
Veröffentlicht: (2024)
Enhancing Motion Variation in Text-to-Motion Models via Pose and Video Conditioned Editing
von: Leite, Clayton, et al.
Veröffentlicht: (2024)
von: Leite, Clayton, et al.
Veröffentlicht: (2024)
Flow Straight and Fast in Hilbert Space: Functional Rectified Flow
von: Zhang, Jianxin, et al.
Veröffentlicht: (2025)
von: Zhang, Jianxin, et al.
Veröffentlicht: (2025)
Label Embedding via Low-Coherence Matrices
von: Zhang, Jianxin, et al.
Veröffentlicht: (2023)
von: Zhang, Jianxin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Transformers, parallel computation, and logarithmic depth
von: Sanford, Clayton, et al.
Veröffentlicht: (2024) -
On Achieving Optimal Adversarial Test Error
von: Li, Justin D., et al.
Veröffentlicht: (2023) -
Astral Space: Convex Analysis at Infinity
von: Dudík, Miroslav, et al.
Veröffentlicht: (2022) -
Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic Regression
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025) -
Spectrum Extraction and Clipping for Implicitly Linear Layers
von: Boroojeny, Ali Ebrahimpour, et al.
Veröffentlicht: (2024)