Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Demir, Samet, Dogan, Zafer |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
von: Demir, Samet, et al.
Veröffentlicht: (2025)
von: Demir, Samet, et al.
Veröffentlicht: (2025)
Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure
von: Demir, Samet, et al.
Veröffentlicht: (2025)
von: Demir, Samet, et al.
Veröffentlicht: (2025)
Input-Label Correlation Governs a Linear-to-Nonlinear Transition in Random Features under Spiked Covariance
von: Demir, Samet, et al.
Veröffentlicht: (2024)
von: Demir, Samet, et al.
Veröffentlicht: (2024)
Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
von: Demir, Samet, et al.
Veröffentlicht: (2025)
von: Demir, Samet, et al.
Veröffentlicht: (2025)
Implicitly Normalized Online PCA: A Regularized Algorithm with Exact High-Dimensional Dynamics
von: Demir, Samet, et al.
Veröffentlicht: (2025)
von: Demir, Samet, et al.
Veröffentlicht: (2025)
Learning Beyond the Gaussian Data: Learning Dynamics of Neural Networks on an Expressive and Cumulant-Controllable Data Model
von: Ure, Onat, et al.
Veröffentlicht: (2026)
von: Ure, Onat, et al.
Veröffentlicht: (2026)
Benefits of Online Tilted Empirical Risk Minimization: A Case Study of Outlier Detection and Robust Regression
von: Yildirim, Yigit E., et al.
Veröffentlicht: (2025)
von: Yildirim, Yigit E., et al.
Veröffentlicht: (2025)
Learning Rate Should Scale Inversely with High-Order Data Moments in High-Dimensional Online Independent Component Analysis
von: Gultekin, M. Oguzhan, et al.
Veröffentlicht: (2025)
von: Gultekin, M. Oguzhan, et al.
Veröffentlicht: (2025)
Learnability and Competition in High-Dimensional Multi-Component ICA
von: Genc, Eser Ilke, et al.
Veröffentlicht: (2026)
von: Genc, Eser Ilke, et al.
Veröffentlicht: (2026)
Exploring the Precise Dynamics of Single-Layer GAN Models: Leveraging Multi-Feature Discriminators for High-Dimensional Subspace Learning
von: Bond, Andrew, et al.
Veröffentlicht: (2024)
von: Bond, Andrew, et al.
Veröffentlicht: (2024)
Test-Time Training Provably Improves Transformers as In-context Learners
von: Gozeten, Halil Alperen, et al.
Veröffentlicht: (2025)
von: Gozeten, Halil Alperen, et al.
Veröffentlicht: (2025)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
Latent Chain-of-Thought Improves Structured-Data Transformers
von: Dudley, Carson, et al.
Veröffentlicht: (2026)
von: Dudley, Carson, et al.
Veröffentlicht: (2026)
Negotiated Representations to Prevent Overfitting in Machine Learning Applications
von: Korhan, Nuri, et al.
Veröffentlicht: (2023)
von: Korhan, Nuri, et al.
Veröffentlicht: (2023)
Asymptotic Learning Curves for Diffusion Models with Random Features Score and Manifold Data
von: George, Anand Jerry, et al.
Veröffentlicht: (2026)
von: George, Anand Jerry, et al.
Veröffentlicht: (2026)
Towards Generalized Hydrological Forecasting using Transformer Models for 120-Hour Streamflow Prediction
von: Demiray, Bekir Z., et al.
Veröffentlicht: (2024)
von: Demiray, Bekir Z., et al.
Veröffentlicht: (2024)
Selective Attention: Enhancing Transformer through Principled Context Control
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
Can Transformers Learn Optimal Filtering for Unknown Systems?
von: Balim, Haldun, et al.
Veröffentlicht: (2023)
von: Balim, Haldun, et al.
Veröffentlicht: (2023)
Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
von: Hayase, Tomohiro, et al.
Veröffentlicht: (2025)
von: Hayase, Tomohiro, et al.
Veröffentlicht: (2025)
Covariance-Aware Transformers for Quadratic Programming and Decision Making
von: Tire, Kutay, et al.
Veröffentlicht: (2026)
von: Tire, Kutay, et al.
Veröffentlicht: (2026)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
All Random Features Representations are Equivalent
von: Sernau, Luke, et al.
Veröffentlicht: (2024)
von: Sernau, Luke, et al.
Veröffentlicht: (2024)
Learning Randomized Algorithms with Transformers
von: von Oswald, Johannes, et al.
Veröffentlicht: (2024)
von: von Oswald, Johannes, et al.
Veröffentlicht: (2024)
Equivalence of Context and Parameter Updates in Modern Transformer Blocks
von: Goldwaser, Adrian, et al.
Veröffentlicht: (2025)
von: Goldwaser, Adrian, et al.
Veröffentlicht: (2025)
Circuit Transformer: A Transformer That Preserves Logical Equivalence
von: Li, Xihan, et al.
Veröffentlicht: (2024)
von: Li, Xihan, et al.
Veröffentlicht: (2024)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
von: Park, Jongho, et al.
Veröffentlicht: (2024)
von: Park, Jongho, et al.
Veröffentlicht: (2024)
Asymptotic theory of in-context learning by linear attention
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
Towards Understanding How Transformers Learn In-context Through a Representation Learning Lens
von: Ren, Ruifeng, et al.
Veröffentlicht: (2023)
von: Ren, Ruifeng, et al.
Veröffentlicht: (2023)
Linking In-context Learning in Transformers to Human Episodic Memory
von: Ji-An, Li, et al.
Veröffentlicht: (2024)
von: Ji-An, Li, et al.
Veröffentlicht: (2024)
Learning with Subset Stacking
von: Birbil, Ş. İlker, et al.
Veröffentlicht: (2021)
von: Birbil, Ş. İlker, et al.
Veröffentlicht: (2021)
Towards Understanding Transformers in Learning Random Walks
von: Shi, Wei, et al.
Veröffentlicht: (2025)
von: Shi, Wei, et al.
Veröffentlicht: (2025)
Learning to Bet for Horizon-Aware Anytime-Valid Testing
von: Taga, Ege Onur, et al.
Veröffentlicht: (2026)
von: Taga, Ege Onur, et al.
Veröffentlicht: (2026)
Cross-Embodied Affordance Transfer through Learning Affordance Equivalences
von: Aktas, Hakan, et al.
Veröffentlicht: (2024)
von: Aktas, Hakan, et al.
Veröffentlicht: (2024)
Asymptotics of Learning with Deep Structured (Random) Features
von: Schröder, Dominik, et al.
Veröffentlicht: (2024)
von: Schröder, Dominik, et al.
Veröffentlicht: (2024)
On the Power of Convolution Augmented Transformer
von: Li, Mingchen, et al.
Veröffentlicht: (2024)
von: Li, Mingchen, et al.
Veröffentlicht: (2024)
Breaking through the learning plateaus of in-context learning in Transformer
von: Fu, Jingwen, et al.
Veröffentlicht: (2023)
von: Fu, Jingwen, et al.
Veröffentlicht: (2023)
CBMAP: Clustering-based manifold approximation and projection for dimensionality reduction
von: Dogan, Berat
Veröffentlicht: (2024)
von: Dogan, Berat
Veröffentlicht: (2024)
Cluster Purge Loss: Structuring Transformer Embeddings for Equivalent Mutants Detection
von: Danilov, Adelaide, et al.
Veröffentlicht: (2025)
von: Danilov, Adelaide, et al.
Veröffentlicht: (2025)
Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation
von: Shi, Kexuan, et al.
Veröffentlicht: (2026)
von: Shi, Kexuan, et al.
Veröffentlicht: (2026)
Transformers are Universal In-context Learners
von: Furuya, Takashi, et al.
Veröffentlicht: (2024)
von: Furuya, Takashi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
von: Demir, Samet, et al.
Veröffentlicht: (2025) -
Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure
von: Demir, Samet, et al.
Veröffentlicht: (2025) -
Input-Label Correlation Governs a Linear-to-Nonlinear Transition in Random Features under Spiked Covariance
von: Demir, Samet, et al.
Veröffentlicht: (2024) -
Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
von: Demir, Samet, et al.
Veröffentlicht: (2025) -
Implicitly Normalized Online PCA: A Regularized Algorithm with Exact High-Dimensional Dynamics
von: Demir, Samet, et al.
Veröffentlicht: (2025)