Anti Mode-Collapse in Mean-Field Transformer via Auxiliary Variables
Fuente:
arXiv
Saved in:
| Main Authors: | Imaizumi, Masaaki, Koyama, Masanori, Isobe, Noboru, Hayashi, Kohei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training-Induced Escape from Token Clustering in a Mean-Field Formulation of Transformers
by: Isobe, Noboru, et al.
Published: (2026)
by: Isobe, Noboru, et al.
Published: (2026)
Minimax Optimal Estimation of Transport-Growth Pairs in Unbalanced Optimal Transport
by: Ponnoprat, Donlapark, et al.
Published: (2026)
by: Ponnoprat, Donlapark, et al.
Published: (2026)
Extended Flow Matching: a Method of Conditional Generation with Generalized Continuity Equation
by: Isobe, Noboru, et al.
Published: (2024)
by: Isobe, Noboru, et al.
Published: (2024)
Neural Fourier Transform: A General Approach to Equivariant Representation Learning
by: Koyama, Masanori, et al.
Published: (2023)
by: Koyama, Masanori, et al.
Published: (2023)
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2026)
by: Nishiyama, Sota, et al.
Published: (2026)
Flow matching achieves almost minimax optimal convergence
by: Fukumizu, Kenji, et al.
Published: (2024)
by: Fukumizu, Kenji, et al.
Published: (2024)
Precise Dynamics of Diagonal Linear Networks: A Unifying Analysis by Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2025)
by: Nishiyama, Sota, et al.
Published: (2025)
Inter-environmental world modeling for continuous and compositional dynamics
by: Hayashi, Kohei, et al.
Published: (2025)
by: Hayashi, Kohei, et al.
Published: (2025)
Spectrum-Adaptive Generalization Bounds for Trained Deep Transformers
by: Sakai, Mana, et al.
Published: (2026)
by: Sakai, Mana, et al.
Published: (2026)
Approximation of Permutation Invariant Polynomials by Transformers: Efficient Construction in Column-Size
by: Takeshita, Naoki, et al.
Published: (2025)
by: Takeshita, Naoki, et al.
Published: (2025)
Pairwise Optimal Transports for Training All-to-All Flow-Based Condition Transfer Model
by: Ikeda, Kotaro, et al.
Published: (2025)
by: Ikeda, Kotaro, et al.
Published: (2025)
Mode Collapse of Mean-Field Variational Inference
by: Sheng, Shunan, et al.
Published: (2025)
by: Sheng, Shunan, et al.
Published: (2025)
Automatic Domain Adaptation by Transformers in In-Context Learning
by: Hataya, Ryuichiro, et al.
Published: (2024)
by: Hataya, Ryuichiro, et al.
Published: (2024)
Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
by: Chen, Baiyuan, et al.
Published: (2025)
by: Chen, Baiyuan, et al.
Published: (2025)
Bayesian Analysis for Over-parameterized Linear Model via Effective Spectra
by: Wakayama, Tomoya, et al.
Published: (2023)
by: Wakayama, Tomoya, et al.
Published: (2023)
Neuron Block Dynamics for XOR Classification with Zero-Margin
by: Braun, Guillaume, et al.
Published: (2026)
by: Braun, Guillaume, et al.
Published: (2026)
High-dimensional Contextual Bandit Problem without Sparsity
by: Komiyama, Junpei, et al.
Published: (2023)
by: Komiyama, Junpei, et al.
Published: (2023)
Zero Generalization Error Theorem for Random Interpolators via Algebraic Geometry
by: Yoshida, Naoki, et al.
Published: (2025)
by: Yoshida, Naoki, et al.
Published: (2025)
Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor Programs
by: Sakai, Mana, et al.
Published: (2025)
by: Sakai, Mana, et al.
Published: (2025)
A convergence result of a continuous model of deep learning via Łojasiewicz--Simon inequality
by: Isobe, Noboru
Published: (2023)
by: Isobe, Noboru
Published: (2023)
Effect of Random Learning Rate: Theoretical Analysis of SGD Dynamics in Non-Convex Optimization via Stationary Distribution
by: Yoshida, Naoki, et al.
Published: (2024)
by: Yoshida, Naoki, et al.
Published: (2024)
Extended Wasserstein-GAN Approach to Causal Distribution Learning: Density-Free Estimation and Minimax Optimality
by: Tamano, Shu, et al.
Published: (2026)
by: Tamano, Shu, et al.
Published: (2026)
Benign Overfitting in Time Series Linear Models with Over-Parameterization
by: Nakakita, Shogo, et al.
Published: (2022)
by: Nakakita, Shogo, et al.
Published: (2022)
Thinking While Listening: Fast-Slow Recurrence for Long-Horizon Sequential Modeling
by: Takashiro, Shota, et al.
Published: (2026)
by: Takashiro, Shota, et al.
Published: (2026)
C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions
by: Kubo, Kenji, et al.
Published: (2026)
by: Kubo, Kenji, et al.
Published: (2026)
Finite-Sample Inference for Sparsely Permuted Linear Regression
by: Ota, Hirofumi, et al.
Published: (2026)
by: Ota, Hirofumi, et al.
Published: (2026)
Loss-Guided Auxiliary Agents for Overcoming Mode Collapse in GFlowNets
by: Malek, Idriss, et al.
Published: (2025)
by: Malek, Idriss, et al.
Published: (2025)
Minimax Rates of Estimation for Optimal Transport Map between Infinite-Dimensional Spaces
by: Ponnoprat, Donlapark, et al.
Published: (2025)
by: Ponnoprat, Donlapark, et al.
Published: (2025)
Sinkhorn Algorithm for Sequentially Composed Optimal Transports
by: Watanabe, Kazuki, et al.
Published: (2024)
by: Watanabe, Kazuki, et al.
Published: (2024)
Effect of Weight Quantization on Learning Models by Typical Case Analysis
by: Kashiwamura, Shuhei, et al.
Published: (2024)
by: Kashiwamura, Shuhei, et al.
Published: (2024)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
High-dimensional Nonparametric Contextual Bandit Problem
by: Iwazaki, Shogo, et al.
Published: (2025)
by: Iwazaki, Shogo, et al.
Published: (2025)
Learning a Single Index Model from Anisotropic Data with vanilla Stochastic Gradient Descent
by: Braun, Guillaume, et al.
Published: (2025)
by: Braun, Guillaume, et al.
Published: (2025)
DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
by: Shing, Makoto, et al.
Published: (2025)
by: Shing, Makoto, et al.
Published: (2025)
Spectral Gradient Descent Mitigates Anisotropy-Driven Misalignment: A Case Study in Phase Retrieval
by: Braun, Guillaume, et al.
Published: (2026)
by: Braun, Guillaume, et al.
Published: (2026)
Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient Descent
by: Imai, Shota, et al.
Published: (2026)
by: Imai, Shota, et al.
Published: (2026)
Fast Escape, Slow Convergence: Learning Dynamics of Phase Retrieval under Power-Law Data
by: Braun, Guillaume, et al.
Published: (2025)
by: Braun, Guillaume, et al.
Published: (2025)
Variational formulations of ODE-Net as a mean-field optimal control problem and existence results
by: Isobe, Noboru, et al.
Published: (2023)
by: Isobe, Noboru, et al.
Published: (2023)
Federated Learning with Relative Fairness
by: Nakakita, Shogo, et al.
Published: (2024)
by: Nakakita, Shogo, et al.
Published: (2024)
Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion
by: Hayakawa, Satoshi, et al.
Published: (2025)
by: Hayakawa, Satoshi, et al.
Published: (2025)
Similar Items
-
Training-Induced Escape from Token Clustering in a Mean-Field Formulation of Transformers
by: Isobe, Noboru, et al.
Published: (2026) -
Minimax Optimal Estimation of Transport-Growth Pairs in Unbalanced Optimal Transport
by: Ponnoprat, Donlapark, et al.
Published: (2026) -
Extended Flow Matching: a Method of Conditional Generation with Generalized Continuity Equation
by: Isobe, Noboru, et al.
Published: (2024) -
Neural Fourier Transform: A General Approach to Equivariant Representation Learning
by: Koyama, Masanori, et al.
Published: (2023) -
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2026)