Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble
Fuente:
arXiv
Guardado en:
| Autores principales: | Nitanda, Atsushi, Lee, Anzelle, Kai, Damian Tan Xing, Sakaguchi, Mizuki, Suzuki, Taiji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improved Particle Approximation Error for Mean Field Neural Networks
por: Nitanda, Atsushi
Publicado: (2024)
por: Nitanda, Atsushi
Publicado: (2024)
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
por: Fu, Guoji, et al.
Publicado: (2026)
por: Fu, Guoji, et al.
Publicado: (2026)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
por: Kawata, Ryotaro, et al.
Publicado: (2025)
por: Kawata, Ryotaro, et al.
Publicado: (2025)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
por: Chen, Zonghao, et al.
Publicado: (2025)
por: Chen, Zonghao, et al.
Publicado: (2025)
Slowly Annealed Langevin Dynamics: Theory and Applications to Training-Free Guided Generation
por: Nitanda, Atsushi, et al.
Publicado: (2026)
por: Nitanda, Atsushi, et al.
Publicado: (2026)
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
por: Kim, Juno, et al.
Publicado: (2023)
por: Kim, Juno, et al.
Publicado: (2023)
Koopman-based generalization bound: New aspect for full-rank weights
por: Hashimoto, Yuka, et al.
Publicado: (2023)
por: Hashimoto, Yuka, et al.
Publicado: (2023)
Uniform convergence of the smooth calibration error and its relationship with functional gradient
por: Futami, Futoshi, et al.
Publicado: (2025)
por: Futami, Futoshi, et al.
Publicado: (2025)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
por: Kim, Juno, et al.
Publicado: (2024)
por: Kim, Juno, et al.
Publicado: (2024)
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
por: Bu, Dake, et al.
Publicado: (2026)
por: Bu, Dake, et al.
Publicado: (2026)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
por: Bu, Dake, et al.
Publicado: (2025)
por: Bu, Dake, et al.
Publicado: (2025)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
por: Bu, Dake, et al.
Publicado: (2024)
por: Bu, Dake, et al.
Publicado: (2024)
Thinned Mean Field Langevin Dynamics
por: Chen, Zonghao, et al.
Publicado: (2026)
por: Chen, Zonghao, et al.
Publicado: (2026)
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
por: Yoshida, Kotaro, et al.
Publicado: (2025)
por: Yoshida, Kotaro, et al.
Publicado: (2025)
Alternating Diffusion for Proximal Sampling with Zeroth Order Queries
por: Takagi, Hirohane, et al.
Publicado: (2026)
por: Takagi, Hirohane, et al.
Publicado: (2026)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
por: Takakura, Shokichi, et al.
Publicado: (2024)
por: Takakura, Shokichi, et al.
Publicado: (2024)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
por: Nishikawa, Naoki, et al.
Publicado: (2024)
por: Nishikawa, Naoki, et al.
Publicado: (2024)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
por: Bu, Dake, et al.
Publicado: (2025)
por: Bu, Dake, et al.
Publicado: (2025)
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
por: Bu, Dake, et al.
Publicado: (2025)
por: Bu, Dake, et al.
Publicado: (2025)
Mirror Mean-Field Langevin Dynamics
por: Gu, Anming, et al.
Publicado: (2025)
por: Gu, Anming, et al.
Publicado: (2025)
Convergence Error Analysis of Reflected Gradient Langevin Dynamics for Globally Optimizing Non-Convex Constrained Problems
por: Sato, Kanji, et al.
Publicado: (2022)
por: Sato, Kanji, et al.
Publicado: (2022)
Beyond Propagation of Chaos: A Stochastic Algorithm for Mean Field Optimization
por: Tankala, Chandan, et al.
Publicado: (2025)
por: Tankala, Chandan, et al.
Publicado: (2025)
Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation
por: Maeda, Ibuki, et al.
Publicado: (2025)
por: Maeda, Ibuki, et al.
Publicado: (2025)
Mirror Descent Policy Optimisation for Robust Constrained Markov Decision Processes
por: Bossens, David M., et al.
Publicado: (2025)
por: Bossens, David M., et al.
Publicado: (2025)
Learning Multi-Index Models with Neural Networks via Mean-Field Langevin Dynamics
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2024)
por: Mousavi-Hosseini, Alireza, et al.
Publicado: (2024)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
por: Higuchi, Rei, et al.
Publicado: (2025)
por: Higuchi, Rei, et al.
Publicado: (2025)
Test time training enhances in-context learning of nonlinear functions
por: Kuwataka, Kento, et al.
Publicado: (2025)
por: Kuwataka, Kento, et al.
Publicado: (2025)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
por: Wakayama, Tomoya, et al.
Publicado: (2025)
por: Wakayama, Tomoya, et al.
Publicado: (2025)
Deep Two-Way Matrix Reordering for Relational Data Analysis
por: Watanabe, Chihiro, et al.
Publicado: (2021)
por: Watanabe, Chihiro, et al.
Publicado: (2021)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
por: Awano, Ryoya, et al.
Publicado: (2026)
por: Awano, Ryoya, et al.
Publicado: (2026)
Transformers Provably Solve Parity Efficiently with Chain of Thought
por: Kim, Juno, et al.
Publicado: (2024)
por: Kim, Juno, et al.
Publicado: (2024)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
por: Watanabe, Chihiro, et al.
Publicado: (2021)
por: Watanabe, Chihiro, et al.
Publicado: (2021)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
por: Kawata, Ryotaro, et al.
Publicado: (2026)
por: Kawata, Ryotaro, et al.
Publicado: (2026)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
por: Takakura, Shokichi, et al.
Publicado: (2023)
por: Takakura, Shokichi, et al.
Publicado: (2023)
Why is parameter averaging beneficial in SGD? An objective smoothing perspective
por: Nitanda, Atsushi, et al.
Publicado: (2023)
por: Nitanda, Atsushi, et al.
Publicado: (2023)
Private Continuous-Time Synthetic Trajectory Generation via Mean-Field Langevin Dynamics
por: Gu, Anming, et al.
Publicado: (2025)
por: Gu, Anming, et al.
Publicado: (2025)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
por: Kim, Juno, et al.
Publicado: (2025)
por: Kim, Juno, et al.
Publicado: (2025)
Mean-Field Langevin Dynamics for Signed Measures via a Bilevel Approach
por: Wang, Guillaume, et al.
Publicado: (2024)
por: Wang, Guillaume, et al.
Publicado: (2024)
Multi-Model Ensemble and Reservoir Computing for River Discharge Prediction in Ungauged Basins
por: Funato, Mizuki, et al.
Publicado: (2025)
por: Funato, Mizuki, et al.
Publicado: (2025)
Improved high-dimensional estimation with Langevin dynamics and stochastic weight averaging
por: Wei, Stanley, et al.
Publicado: (2026)
por: Wei, Stanley, et al.
Publicado: (2026)
Ejemplares similares
-
Improved Particle Approximation Error for Mean Field Neural Networks
por: Nitanda, Atsushi
Publicado: (2024) -
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
por: Fu, Guoji, et al.
Publicado: (2026) -
Direct Distributional Optimization for Provable Alignment of Diffusion Models
por: Kawata, Ryotaro, et al.
Publicado: (2025) -
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
por: Chen, Zonghao, et al.
Publicado: (2025) -
Slowly Annealed Langevin Dynamics: Theory and Applications to Training-Free Guided Generation
por: Nitanda, Atsushi, et al.
Publicado: (2026)