How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
Fuente:
arXiv
Saved in:
| Main Authors: | Demir, Samet, Dogan, Zafer |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
by: Demir, Samet, et al.
Published: (2025)
by: Demir, Samet, et al.
Published: (2025)
Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure
by: Demir, Samet, et al.
Published: (2025)
by: Demir, Samet, et al.
Published: (2025)
Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
by: Demir, Samet, et al.
Published: (2025)
by: Demir, Samet, et al.
Published: (2025)
Learning Beyond the Gaussian Data: Learning Dynamics of Neural Networks on an Expressive and Cumulant-Controllable Data Model
by: Ure, Onat, et al.
Published: (2026)
by: Ure, Onat, et al.
Published: (2026)
Implicitly Normalized Online PCA: A Regularized Algorithm with Exact High-Dimensional Dynamics
by: Demir, Samet, et al.
Published: (2025)
by: Demir, Samet, et al.
Published: (2025)
Input-Label Correlation Governs a Linear-to-Nonlinear Transition in Random Features under Spiked Covariance
by: Demir, Samet, et al.
Published: (2024)
by: Demir, Samet, et al.
Published: (2024)
Learning Rate Should Scale Inversely with High-Order Data Moments in High-Dimensional Online Independent Component Analysis
by: Gultekin, M. Oguzhan, et al.
Published: (2025)
by: Gultekin, M. Oguzhan, et al.
Published: (2025)
Benefits of Online Tilted Empirical Risk Minimization: A Case Study of Outlier Detection and Robust Regression
by: Yildirim, Yigit E., et al.
Published: (2025)
by: Yildirim, Yigit E., et al.
Published: (2025)
Learnability and Competition in High-Dimensional Multi-Component ICA
by: Genc, Eser Ilke, et al.
Published: (2026)
by: Genc, Eser Ilke, et al.
Published: (2026)
Exploring the Precise Dynamics of Single-Layer GAN Models: Leveraging Multi-Feature Discriminators for High-Dimensional Subspace Learning
by: Bond, Andrew, et al.
Published: (2024)
by: Bond, Andrew, et al.
Published: (2024)
MLPs Learn In-Context on Regression and Classification Tasks
by: Tong, William L., et al.
Published: (2024)
by: Tong, William L., et al.
Published: (2024)
Latent Chain-of-Thought Improves Structured-Data Transformers
by: Dudley, Carson, et al.
Published: (2026)
by: Dudley, Carson, et al.
Published: (2026)
When and How Unlabeled Data Provably Improve In-Context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
Bayesian Inference with Shaped Deep Non-linear MLPs
by: Hanin, Boris, et al.
Published: (2026)
by: Hanin, Boris, et al.
Published: (2026)
Constructing Efficient Fact-Storing MLPs for Transformers
by: Dugan, Owen, et al.
Published: (2025)
by: Dugan, Owen, et al.
Published: (2025)
Is In-Context Universality Enough? MLPs are Also Universal In-Context
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
by: Park, Jongho, et al.
Published: (2024)
by: Park, Jongho, et al.
Published: (2024)
Benchmarking Optimizers for MLPs in Tabular Deep Learning
by: Gorishniy, Yury, et al.
Published: (2026)
by: Gorishniy, Yury, et al.
Published: (2026)
Equivalence of Context and Parameter Updates in Modern Transformer Blocks
by: Goldwaser, Adrian, et al.
Published: (2025)
by: Goldwaser, Adrian, et al.
Published: (2025)
Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
by: Chowdhury, Tanya, et al.
Published: (2025)
by: Chowdhury, Tanya, et al.
Published: (2025)
Why Attention Fails: The Degeneration of Transformers into MLPs in Time Series Forecasting
by: Liang, Zida, et al.
Published: (2025)
by: Liang, Zida, et al.
Published: (2025)
In-Context Learning Under Regime Change
by: Dudley, Carson, et al.
Published: (2026)
by: Dudley, Carson, et al.
Published: (2026)
Converting MLPs into Polynomials in Closed Form
by: Belrose, Nora, et al.
Published: (2025)
by: Belrose, Nora, et al.
Published: (2025)
Interpolated-MLPs: Controllable Inductive Bias
by: Wu, Sean, et al.
Published: (2024)
by: Wu, Sean, et al.
Published: (2024)
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
How do Transformers perform In-Context Autoregressive Learning?
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
Better by Default: Strong Pre-Tuned MLPs and Boosted Trees on Tabular Data
by: Holzmüller, David, et al.
Published: (2024)
by: Holzmüller, David, et al.
Published: (2024)
MLPs at the EOC: Spectrum of the NTK
by: Terjék, Dávid, et al.
Published: (2025)
by: Terjék, Dávid, et al.
Published: (2025)
MLPs at the EOC: Concentration of the NTK
by: Terjék, Dávid, et al.
Published: (2025)
by: Terjék, Dávid, et al.
Published: (2025)
Selective Attention: Enhancing Transformer through Principled Context Control
by: Zhang, Xuechen, et al.
Published: (2024)
by: Zhang, Xuechen, et al.
Published: (2024)
From Latent Space to Training Data: Explainable Specialization in Minimal MLPs
by: Alba, Enrique, et al.
Published: (2026)
by: Alba, Enrique, et al.
Published: (2026)
TransformMix: Learning Transformation and Mixing Strategies from Data
by: Cheung, Tsz-Him, et al.
Published: (2024)
by: Cheung, Tsz-Him, et al.
Published: (2024)
Dimensional Criticality at Grokking Across MLPs and Transformers
by: Wang, Ping
Published: (2026)
by: Wang, Ping
Published: (2026)
From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
by: Kawata, Ryotaro, et al.
Published: (2025)
by: Kawata, Ryotaro, et al.
Published: (2025)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
by: Kahardipraja, Patrick, et al.
Published: (2025)
by: Kahardipraja, Patrick, et al.
Published: (2025)
Bilinear MLPs enable weight-based mechanistic interpretability
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
Hyperparameter Tuning MLPs for Probabilistic Time Series Forecasting
by: Madhusudhanan, Kiran, et al.
Published: (2024)
by: Madhusudhanan, Kiran, et al.
Published: (2024)
The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
by: Pengmei, Zihan, et al.
Published: (2025)
by: Pengmei, Zihan, et al.
Published: (2025)
Negotiated Representations to Prevent Overfitting in Machine Learning Applications
by: Korhan, Nuri, et al.
Published: (2023)
by: Korhan, Nuri, et al.
Published: (2023)
Mimetic Initialization of MLPs
by: Trockman, Asher, et al.
Published: (2026)
by: Trockman, Asher, et al.
Published: (2026)
Similar Items
-
Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
by: Demir, Samet, et al.
Published: (2025) -
Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure
by: Demir, Samet, et al.
Published: (2025) -
Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
by: Demir, Samet, et al.
Published: (2025) -
Learning Beyond the Gaussian Data: Learning Dynamics of Neural Networks on an Expressive and Cumulant-Controllable Data Model
by: Ure, Onat, et al.
Published: (2026) -
Implicitly Normalized Online PCA: A Regularized Algorithm with Exact High-Dimensional Dynamics
by: Demir, Samet, et al.
Published: (2025)