How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongkang, Wang, Meng, Lu, Songtao, Cui, Xiaodong, Chen, Pin-Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2025)
by: Li, Hongkang, et al.
Published: (2025)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
by: Li, Hongkang, et al.
Published: (2025)
by: Li, Hongkang, et al.
Published: (2025)
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
by: Song, Bingqing, et al.
Published: (2025)
by: Song, Bingqing, et al.
Published: (2025)
How does promoting the minority fraction affect generalization? A theoretical study of the one-hidden-layer neural network on group imbalance
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Theoretical Learning Performance of Graph Neural Networks: The Impact of Jumping Connections and Layer-wise Sparsification
by: Sun, Jiawei, et al.
Published: (2025)
by: Sun, Jiawei, et al.
Published: (2025)
Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs
by: Li, Hongkang, et al.
Published: (2026)
by: Li, Hongkang, et al.
Published: (2026)
SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning
by: Zhang, Shuai, et al.
Published: (2024)
by: Zhang, Shuai, et al.
Published: (2024)
Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study
by: Zhang, Xingxuan, et al.
Published: (2025)
by: Zhang, Xingxuan, et al.
Published: (2025)
Provable In-Context Learning of Nonlinear Regression with Transformers
by: Li, Hongbo, et al.
Published: (2025)
by: Li, Hongbo, et al.
Published: (2025)
Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
by: Ding, Yanna, et al.
Published: (2025)
by: Ding, Yanna, et al.
Published: (2025)
How do Transformers perform In-Context Autoregressive Learning?
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
Understanding Generalization and Forgetting in In-Context Continual Learning
by: Li, Guangyu, et al.
Published: (2026)
by: Li, Guangyu, et al.
Published: (2026)
A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models
by: Shandirasegaran, Mugunthan, et al.
Published: (2026)
by: Shandirasegaran, Mugunthan, et al.
Published: (2026)
On the Role of Transformer Feed-Forward Layers in Nonlinear In-Context Learning
by: Sun, Haoyuan, et al.
Published: (2025)
by: Sun, Haoyuan, et al.
Published: (2025)
Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer
by: Hsu, Alexander, et al.
Published: (2026)
by: Hsu, Alexander, et al.
Published: (2026)
How Transformers Learn In-Context Recall Tasks? Optimality, Training Dynamics and Generalization
by: Nguyen, Quan, et al.
Published: (2025)
by: Nguyen, Quan, et al.
Published: (2025)
In-Context Compositional Learning via Sparse Coding Transformer
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
Visual prompting reimagined: The power of the Activation Prompts
by: Zhang, Yihua, et al.
Published: (2026)
by: Zhang, Yihua, et al.
Published: (2026)
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Transformers Meet In-Context Learning: A Universal Approximation Theory
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Generative Pre-Trained Transformer for Symbolic Regression Base In-Context Reinforcement Learning
by: Li, Yanjie, et al.
Published: (2024)
by: Li, Yanjie, et al.
Published: (2024)
How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
by: Demir, Samet, et al.
Published: (2025)
by: Demir, Samet, et al.
Published: (2025)
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization
by: Saif, A F M, et al.
Published: (2024)
by: Saif, A F M, et al.
Published: (2024)
Meta-Learning Transformers to Improve In-Context Generalization
by: Braccaioli, Lorenzo, et al.
Published: (2025)
by: Braccaioli, Lorenzo, et al.
Published: (2025)
In-Context In-Context Learning with Transformer Neural Processes
by: Ashman, Matthew, et al.
Published: (2024)
by: Ashman, Matthew, et al.
Published: (2024)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
by: Chen, Xingwu, et al.
Published: (2024)
by: Chen, Xingwu, et al.
Published: (2024)
Computational Safety for Generative AI: A Signal Processing Perspective
by: Chen, Pin-Yu
Published: (2025)
by: Chen, Pin-Yu
Published: (2025)
Node Identifiers: Compact, Discrete Representations for Efficient Graph Learning
by: Luo, Yuankai, et al.
Published: (2024)
by: Luo, Yuankai, et al.
Published: (2024)
How Do Transformers Learn Variable Binding in Symbolic Programs?
by: Wu, Yiwei, et al.
Published: (2025)
by: Wu, Yiwei, et al.
Published: (2025)
Transformers Can Learn Temporal Difference Methods for In-Context Reinforcement Learning
by: Wang, Jiuqi, et al.
Published: (2024)
by: Wang, Jiuqi, et al.
Published: (2024)
Improving Transformers using Faithful Positional Encoding
by: Idé, Tsuyoshi, et al.
Published: (2024)
by: Idé, Tsuyoshi, et al.
Published: (2024)
Benchmarking General-Purpose In-Context Learning
by: Wang, Fan, et al.
Published: (2024)
by: Wang, Fan, et al.
Published: (2024)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
by: Im, Shawn, et al.
Published: (2026)
by: Im, Shawn, et al.
Published: (2026)
Learning Mutual Excitation for Hand-to-Hand and Human-to-Human Interaction Recognition
by: Liu, Mengyuan, et al.
Published: (2024)
by: Liu, Mengyuan, et al.
Published: (2024)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
General-Purpose In-Context Learning by Meta-Learning Transformers
by: Kirsch, Louis, et al.
Published: (2022)
by: Kirsch, Louis, et al.
Published: (2022)
Similar Items
-
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2025) -
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024) -
When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
by: Li, Hongkang, et al.
Published: (2025) -
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
by: Li, Hongkang, et al.
Published: (2024) -
What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
by: Li, Hongkang, et al.
Published: (2024)