One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zihao, Cao, Yuan, Gao, Cheng, He, Yihan, Liu, Han, Klusowski, Jason M., Fan, Jianqing, Wang, Mengdi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023)
by: Cattaneo, Matias D., et al.
Published: (2023)
Global Convergence in Training Large-Scale Transformers
by: Gao, Cheng, et al.
Published: (2024)
by: Gao, Cheng, et al.
Published: (2024)
Decoding Game: On Minimax Optimality of Heuristic Text Generation Strategies
by: Chen, Sijin, et al.
Published: (2024)
by: Chen, Sijin, et al.
Published: (2024)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)
by: Sheen, Heejune, et al.
Published: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
When and How Unlabeled Data Provably Improve In-Context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
FMIP: Joint Continuous-Integer Flow For Mixed-Integer Linear Programming
by: Li, Hongpei, et al.
Published: (2025)
by: Li, Hongpei, et al.
Published: (2025)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
by: Shen, Jucheng, et al.
Published: (2026)
by: Shen, Jucheng, et al.
Published: (2026)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
One Rank at a Time: Cascading Error Dynamics in Sequential Learning
by: Vandchali, Mahtab Alizadeh, et al.
Published: (2025)
by: Vandchali, Mahtab Alizadeh, et al.
Published: (2025)
Momentum Benefits Non-IID Federated Learning Simply and Provably
by: Cheng, Ziheng, et al.
Published: (2023)
by: Cheng, Ziheng, et al.
Published: (2023)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Stability of Transformers under Layer Normalization
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
by: Lee, Kyungbok, et al.
Published: (2026)
by: Lee, Kyungbok, et al.
Published: (2026)
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning
by: Zhang, Thomas T., et al.
Published: (2025)
by: Zhang, Thomas T., et al.
Published: (2025)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024)
by: Xiao, Minheng, et al.
Published: (2024)
Towards Simple and Provable Parameter-Free Adaptive Gradient Methods
by: Tao, Yuanzhe, et al.
Published: (2024)
by: Tao, Yuanzhe, et al.
Published: (2024)
Provably Safe Generative Sampling with Constricting Barrier Functions
by: Gadginmath, Darshan, et al.
Published: (2026)
by: Gadginmath, Darshan, et al.
Published: (2026)
GLinSAT: The General Linear Satisfiability Neural Network Layer By Accelerated Gradient Descent
by: Zeng, Hongtai, et al.
Published: (2024)
by: Zeng, Hongtai, et al.
Published: (2024)
Adaptive Stabilization Based on Machine Learning for Column Generation
by: Shen, Yunzhuang, et al.
Published: (2024)
by: Shen, Yunzhuang, et al.
Published: (2024)
Using Laplace Transform To Optimize the Hallucination of Generation Models
by: Kang, Cheng, et al.
Published: (2026)
by: Kang, Cheng, et al.
Published: (2026)
$K$-Nearest-Neighbor Resampling for Off-Policy Evaluation in Stochastic Control
by: Giegrich, Michael, et al.
Published: (2023)
by: Giegrich, Michael, et al.
Published: (2023)
One-shot Learning for MIPs with SOS1 Constraints
by: La Rocca, Charly Robinson, et al.
Published: (2024)
by: La Rocca, Charly Robinson, et al.
Published: (2024)
Provable Acceleration for Diffusion Models under Minimal Assumptions
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
How Well Can Transformers Emulate In-context Newton's Method?
by: Giannou, Angeliki, et al.
Published: (2024)
by: Giannou, Angeliki, et al.
Published: (2024)
SMiLE: Provably Enforcing Global Relational Properties in Neural Networks
by: Francobaldi, Matteo, et al.
Published: (2025)
by: Francobaldi, Matteo, et al.
Published: (2025)
Correlated Noise Provably Beats Independent Noise for Differentially Private Learning
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
Neural-Rendezvous: Provably Robust Guidance and Control to Encounter Interstellar Objects
by: Tsukamoto, Hiroyasu, et al.
Published: (2022)
by: Tsukamoto, Hiroyasu, et al.
Published: (2022)
Online Bidding for Contextual First-Price Auctions with Budgets under One-Sided Information Feedback
by: Fu, Zeng, et al.
Published: (2026)
by: Fu, Zeng, et al.
Published: (2026)
Safety-Aware Reinforcement Learning for Electric Vehicle Charging Station Management in Distribution Network
by: Fan, Jiarong, et al.
Published: (2024)
by: Fan, Jiarong, et al.
Published: (2024)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)
by: Liu, Xin, et al.
Published: (2022)
Agent-Based Decentralized Energy Management of EV Charging Station with Solar Photovoltaics via Multi-Agent Reinforcement Learning
by: Fan, Jiarong, et al.
Published: (2025)
by: Fan, Jiarong, et al.
Published: (2025)
Lyapunov Function Consistent Adaptive Network Signal Control with Back Pressure and Reinforcement Learning
by: Ma, Chaolun, et al.
Published: (2022)
by: Ma, Chaolun, et al.
Published: (2022)
The Nearest Graph Laplacian in Frobenius Norm
by: Sato, Kazuhiro, et al.
Published: (2024)
by: Sato, Kazuhiro, et al.
Published: (2024)
Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent
by: Chu, Ya-Chi, et al.
Published: (2025)
by: Chu, Ya-Chi, et al.
Published: (2025)
Safe Decentralized Operation of EV Virtual Power Plant with Limited Network Visibility via Multi-Agent Reinforcement Learning
by: Huang, Chenghao, et al.
Published: (2026)
by: Huang, Chenghao, et al.
Published: (2026)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
A Deep Q-Network Based on Radial Basis Functions for Multi-Echelon Inventory Management
by: Cheng, Liqiang, et al.
Published: (2024)
by: Cheng, Liqiang, et al.
Published: (2024)
Similar Items
-
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023) -
Global Convergence in Training Large-Scale Transformers
by: Gao, Cheng, et al.
Published: (2024) -
Decoding Game: On Minimax Optimality of Heuristic Text Generation Strategies
by: Chen, Sijin, et al.
Published: (2024) -
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026) -
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)