Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Takakura, Shokichi, Suzuki, Taiji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023)
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
von: Higuchi, Rei, et al.
Veröffentlicht: (2026)
von: Higuchi, Rei, et al.
Veröffentlicht: (2026)
Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
von: Takakura, Shokichi, et al.
Veröffentlicht: (2026)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2026)
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
von: Wachi, Akifumi, et al.
Veröffentlicht: (2026)
von: Wachi, Akifumi, et al.
Veröffentlicht: (2026)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
Accelerating Differentially Private Federated Learning via Adaptive Extrapolation
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
FedDuA: Doubly Adaptive Federated Learning
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
Differentially Private Sampling from Distributions via Wasserstein Projection
von: Takakura, Shokichi, et al.
Veröffentlicht: (2026)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2026)
Deep Two-Way Matrix Reordering for Relational Data Analysis
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
Optimal Variance and Covariance Estimation under Differential Privacy in the Add-Remove Model and Beyond
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2025)
DPSQL+: A Differentially Private SQL Library with a Minimum Frequency Rule
von: Matsumoto, Tomoya, et al.
Veröffentlicht: (2026)
von: Matsumoto, Tomoya, et al.
Veröffentlicht: (2026)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
von: Zhang, Tongcheng, et al.
Veröffentlicht: (2026)
von: Zhang, Tongcheng, et al.
Veröffentlicht: (2026)
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
von: Kim, Juno, et al.
Veröffentlicht: (2023)
von: Kim, Juno, et al.
Veröffentlicht: (2023)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
von: Li, Bingrui, et al.
Veröffentlicht: (2024)
von: Li, Bingrui, et al.
Veröffentlicht: (2024)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
von: Chen, Yihang, et al.
Veröffentlicht: (2024)
von: Chen, Yihang, et al.
Veröffentlicht: (2024)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
von: Awano, Ryoya, et al.
Veröffentlicht: (2026)
von: Awano, Ryoya, et al.
Veröffentlicht: (2026)
Quantifying the Optimization and Generalization Advantages of Graph Neural Networks Over Multilayer Perceptrons
von: Huang, Wei, et al.
Veröffentlicht: (2023)
von: Huang, Wei, et al.
Veröffentlicht: (2023)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
Test time training enhances in-context learning of nonlinear functions
von: Kuwataka, Kento, et al.
Veröffentlicht: (2025)
von: Kuwataka, Kento, et al.
Veröffentlicht: (2025)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025)
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025)
Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
von: Chen, Yilan, et al.
Veröffentlicht: (2025)
von: Chen, Yilan, et al.
Veröffentlicht: (2025)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble
von: Nitanda, Atsushi, et al.
Veröffentlicht: (2025)
von: Nitanda, Atsushi, et al.
Veröffentlicht: (2025)
Training Guarantees of Neural Network Classification Two-Sample Tests by Kernel Analysis
von: Khurana, Varun, et al.
Veröffentlicht: (2024)
von: Khurana, Varun, et al.
Veröffentlicht: (2024)
Neural-Kernel Conditional Mean Embeddings
von: Shimizu, Eiki, et al.
Veröffentlicht: (2024)
von: Shimizu, Eiki, et al.
Veröffentlicht: (2024)
Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
von: Lee, Jason D., et al.
Veröffentlicht: (2024)
von: Lee, Jason D., et al.
Veröffentlicht: (2024)
Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression
von: Kim, Juno, et al.
Veröffentlicht: (2025)
von: Kim, Juno, et al.
Veröffentlicht: (2025)
Transformers are Minimax Optimal Nonparametric In-Context Learners
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2026)
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2026)
Stochastic Gradient Descent for Two-layer Neural Networks
von: Cao, Dinghao, et al.
Veröffentlicht: (2024)
von: Cao, Dinghao, et al.
Veröffentlicht: (2024)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
von: Yamamoto, Naoya, et al.
Veröffentlicht: (2025)
von: Yamamoto, Naoya, et al.
Veröffentlicht: (2025)
Convergence Error Analysis of Reflected Gradient Langevin Dynamics for Globally Optimizing Non-Convex Constrained Problems
von: Sato, Kanji, et al.
Veröffentlicht: (2022)
von: Sato, Kanji, et al.
Veröffentlicht: (2022)
Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations
von: Oko, Kazusato, et al.
Veröffentlicht: (2024)
von: Oko, Kazusato, et al.
Veröffentlicht: (2024)
Pretrained transformer efficiently learns low-dimensional target functions in-context
von: Oko, Kazusato, et al.
Veröffentlicht: (2024)
von: Oko, Kazusato, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023) -
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
von: Higuchi, Rei, et al.
Veröffentlicht: (2026) -
Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
von: Takakura, Shokichi, et al.
Veröffentlicht: (2026) -
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
von: Wachi, Akifumi, et al.
Veröffentlicht: (2026) -
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
von: Kim, Juno, et al.
Veröffentlicht: (2024)