Direct Distributional Optimization for Provable Alignment of Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kawata, Ryotaro, Oko, Kazusato, Nitanda, Atsushi, Suzuki, Taiji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations
by: Oko, Kazusato, et al.
Published: (2024)
by: Oko, Kazusato, et al.
Published: (2024)
Pretrained transformer efficiently learns low-dimensional target functions in-context
by: Oko, Kazusato, et al.
Published: (2024)
by: Oko, Kazusato, et al.
Published: (2024)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
by: Kawata, Ryotaro, et al.
Published: (2026)
by: Kawata, Ryotaro, et al.
Published: (2026)
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
by: Kim, Juno, et al.
Published: (2023)
by: Kim, Juno, et al.
Published: (2023)
Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
by: Lee, Jason D., et al.
Published: (2024)
by: Lee, Jason D., et al.
Published: (2024)
Flow matching achieves almost minimax optimal convergence
by: Fukumizu, Kenji, et al.
Published: (2024)
by: Fukumizu, Kenji, et al.
Published: (2024)
When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
by: Kawata, Ryotaro, et al.
Published: (2025)
by: Kawata, Ryotaro, et al.
Published: (2025)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
by: Chen, Zonghao, et al.
Published: (2025)
by: Chen, Zonghao, et al.
Published: (2025)
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
by: Fu, Guoji, et al.
Published: (2026)
by: Fu, Guoji, et al.
Published: (2026)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
by: Bu, Dake, et al.
Published: (2024)
by: Bu, Dake, et al.
Published: (2024)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
by: Higuchi, Rei, et al.
Published: (2026)
by: Higuchi, Rei, et al.
Published: (2026)
Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble
by: Nitanda, Atsushi, et al.
Published: (2025)
by: Nitanda, Atsushi, et al.
Published: (2025)
Koopman-based generalization bound: New aspect for full-rank weights
by: Hashimoto, Yuka, et al.
Published: (2023)
by: Hashimoto, Yuka, et al.
Published: (2023)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
by: Nishikawa, Naoki, et al.
Published: (2024)
by: Nishikawa, Naoki, et al.
Published: (2024)
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
by: Bu, Dake, et al.
Published: (2026)
by: Bu, Dake, et al.
Published: (2026)
Alternating Diffusion for Proximal Sampling with Zeroth Order Queries
by: Takagi, Hirohane, et al.
Published: (2026)
by: Takagi, Hirohane, et al.
Published: (2026)
Improved Particle Approximation Error for Mean Field Neural Networks
by: Nitanda, Atsushi
Published: (2024)
by: Nitanda, Atsushi
Published: (2024)
A Statistical Theory of Contrastive Pre-training and Multimodal Generative AI
by: Oko, Kazusato, et al.
Published: (2025)
by: Oko, Kazusato, et al.
Published: (2025)
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
by: Wakayama, Tomoya, et al.
Published: (2025)
by: Wakayama, Tomoya, et al.
Published: (2025)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
by: Kawata, Ryotaro, et al.
Published: (2025)
by: Kawata, Ryotaro, et al.
Published: (2025)
Uniform convergence of the smooth calibration error and its relationship with functional gradient
by: Futami, Futoshi, et al.
Published: (2025)
by: Futami, Futoshi, et al.
Published: (2025)
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
by: Yoshida, Kotaro, et al.
Published: (2025)
by: Yoshida, Kotaro, et al.
Published: (2025)
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation
by: Maeda, Ibuki, et al.
Published: (2025)
by: Maeda, Ibuki, et al.
Published: (2025)
Mirror Descent Policy Optimisation for Robust Constrained Markov Decision Processes
by: Bossens, David M., et al.
Published: (2025)
by: Bossens, David M., et al.
Published: (2025)
Test time training enhances in-context learning of nonlinear functions
by: Kuwataka, Kento, et al.
Published: (2025)
by: Kuwataka, Kento, et al.
Published: (2025)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
by: Awano, Ryoya, et al.
Published: (2026)
by: Awano, Ryoya, et al.
Published: (2026)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
by: Takakura, Shokichi, et al.
Published: (2024)
by: Takakura, Shokichi, et al.
Published: (2024)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
by: Takakura, Shokichi, et al.
Published: (2023)
by: Takakura, Shokichi, et al.
Published: (2023)
Why is parameter averaging beneficial in SGD? An objective smoothing perspective
by: Nitanda, Atsushi, et al.
Published: (2023)
by: Nitanda, Atsushi, et al.
Published: (2023)
Similar Items
-
Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations
by: Oko, Kazusato, et al.
Published: (2024) -
Pretrained transformer efficiently learns low-dimensional target functions in-context
by: Oko, Kazusato, et al.
Published: (2024) -
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
by: Kawata, Ryotaro, et al.
Published: (2026) -
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
by: Kim, Juno, et al.
Published: (2023) -
Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
by: Lee, Jason D., et al.
Published: (2024)