Transformers are Minimax Optimal Nonparametric In-Context Learners
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Juno, Nakamaki, Tai, Suzuki, Taiji |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
by: Kawata, Ryotaro, et al.
Published: (2026)
by: Kawata, Ryotaro, et al.
Published: (2026)
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
by: Kim, Juno, et al.
Published: (2023)
by: Kim, Juno, et al.
Published: (2023)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
by: Yamamoto, Naoya, et al.
Published: (2025)
by: Yamamoto, Naoya, et al.
Published: (2025)
Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers
by: Ching, Michelle, et al.
Published: (2026)
by: Ching, Michelle, et al.
Published: (2026)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
by: Wakayama, Tomoya, et al.
Published: (2025)
by: Wakayama, Tomoya, et al.
Published: (2025)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
by: Nishikawa, Naoki, et al.
Published: (2024)
by: Nishikawa, Naoki, et al.
Published: (2024)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
by: Takakura, Shokichi, et al.
Published: (2023)
by: Takakura, Shokichi, et al.
Published: (2023)
Nonparametric Instrumental Regression via Kernel Methods is Minimax Optimal
by: Meunier, Dimitri, et al.
Published: (2024)
by: Meunier, Dimitri, et al.
Published: (2024)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
by: Chen, Zonghao, et al.
Published: (2025)
by: Chen, Zonghao, et al.
Published: (2025)
How do Transformers perform In-Context Autoregressive Learning?
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
by: Nishikawa, Naoki, et al.
Published: (2025)
by: Nishikawa, Naoki, et al.
Published: (2025)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
by: Oh, Junsoo, et al.
Published: (2025)
by: Oh, Junsoo, et al.
Published: (2025)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
by: Takakura, Shokichi, et al.
Published: (2024)
by: Takakura, Shokichi, et al.
Published: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
by: Awano, Ryoya, et al.
Published: (2026)
by: Awano, Ryoya, et al.
Published: (2026)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
Test time training enhances in-context learning of nonlinear functions
by: Kuwataka, Kento, et al.
Published: (2025)
by: Kuwataka, Kento, et al.
Published: (2025)
Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,λ}$ Targets
by: Lai, Yanming, et al.
Published: (2026)
by: Lai, Yanming, et al.
Published: (2026)
Nonparametric Teaching of Attention Learners
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
by: Bu, Dake, et al.
Published: (2024)
by: Bu, Dake, et al.
Published: (2024)
Nonparametric Teaching for Graph Property Learners
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits
by: Cho, Brian, et al.
Published: (2024)
by: Cho, Brian, et al.
Published: (2024)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
by: Jiang, Jiarui, et al.
Published: (2025)
by: Jiang, Jiarui, et al.
Published: (2025)
Minimax-optimal and Locally-adaptive Online Nonparametric Regression
by: Liautaud, Paul, et al.
Published: (2024)
by: Liautaud, Paul, et al.
Published: (2024)
Linear Transformers are Versatile In-Context Learners
by: Vladymyrov, Max, et al.
Published: (2024)
by: Vladymyrov, Max, et al.
Published: (2024)
Minimax Adaptive Online Nonparametric Regression over Besov Spaces
by: Liautaud, Paul, et al.
Published: (2025)
by: Liautaud, Paul, et al.
Published: (2025)
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
by: Lascu, Razvan-Andrei, et al.
Published: (2026)
by: Lascu, Razvan-Andrei, et al.
Published: (2026)
Transfer Learning for Nonparametric Regression: Non-asymptotic Minimax Analysis and Adaptive Procedure
by: Cai, T. Tony, et al.
Published: (2024)
by: Cai, T. Tony, et al.
Published: (2024)
In-Context Learning as Nonparametric Conditional Probability Estimation: Risk Bounds and Optimality
by: Liu, Chenrui, et al.
Published: (2025)
by: Liu, Chenrui, et al.
Published: (2025)
Optimal Transport Barycenter via Nonconvex-Concave Minimax Optimization
by: Kim, Kaheon, et al.
Published: (2025)
by: Kim, Kaheon, et al.
Published: (2025)
Weighted Point Set Embedding for Multimodal Contrastive Learning Toward Optimal Similarity Metric
by: Uesaka, Toshimitsu, et al.
Published: (2024)
by: Uesaka, Toshimitsu, et al.
Published: (2024)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
by: Li, Bingrui, et al.
Published: (2024)
by: Li, Bingrui, et al.
Published: (2024)
Minimax Optimal Reinforcement Learning with Quasi-Optimism
by: Lee, Harin, et al.
Published: (2025)
by: Lee, Harin, et al.
Published: (2025)
Linear Bandits on Ellipsoids: Minimax Optimal Algorithms
by: Zhang, Raymond, et al.
Published: (2025)
by: Zhang, Raymond, et al.
Published: (2025)
Minimax Optimal Q Learning with Nearest Neighbors
by: Zhao, Puning, et al.
Published: (2023)
by: Zhao, Puning, et al.
Published: (2023)
Similar Items
-
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024) -
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024) -
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
by: Kawata, Ryotaro, et al.
Published: (2026) -
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
by: Kim, Juno, et al.
Published: (2023) -
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
by: Yamamoto, Naoya, et al.
Published: (2025)