Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kawata, Ryotaro, Matsutani, Kohsei, Kinoshita, Yuri, Nishikawa, Naoki, Suzuki, Taiji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
Dataset Distillation Efficiently Encodes Low-Dimensional Representations from Gradient-Based Learning of Non-Linear Tasks
von: Kinoshita, Yuri, et al.
Veröffentlicht: (2026)
von: Kinoshita, Yuri, et al.
Veröffentlicht: (2026)
From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025)
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025)
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
von: Higuchi, Rei, et al.
Veröffentlicht: (2026)
von: Higuchi, Rei, et al.
Veröffentlicht: (2026)
When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
Transformers Provably Solve Parity Efficiently with Chain of Thought
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
von: Awano, Ryoya, et al.
Veröffentlicht: (2026)
von: Awano, Ryoya, et al.
Veröffentlicht: (2026)
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
von: Liao, Fangshuo, et al.
Veröffentlicht: (2025)
von: Liao, Fangshuo, et al.
Veröffentlicht: (2025)
Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples
von: Bu, Dake, et al.
Veröffentlicht: (2024)
von: Bu, Dake, et al.
Veröffentlicht: (2024)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
von: Bu, Dake, et al.
Veröffentlicht: (2024)
von: Bu, Dake, et al.
Veröffentlicht: (2024)
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
von: Huang, Wei, et al.
Veröffentlicht: (2026)
von: Huang, Wei, et al.
Veröffentlicht: (2026)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
von: Kim, Juno, et al.
Veröffentlicht: (2025)
von: Kim, Juno, et al.
Veröffentlicht: (2025)
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
von: Wachi, Akifumi, et al.
Veröffentlicht: (2026)
von: Wachi, Akifumi, et al.
Veröffentlicht: (2026)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
von: Yamamoto, Naoya, et al.
Veröffentlicht: (2025)
von: Yamamoto, Naoya, et al.
Veröffentlicht: (2025)
Unraveling the Localized Latents: Learning Stratified Manifold Structures in LLM Embedding Space with Sparse Mixture-of-Experts
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
Weighted Point Set Embedding for Multimodal Contrastive Learning Toward Optimal Similarity Metric
von: Uesaka, Toshimitsu, et al.
Veröffentlicht: (2024)
von: Uesaka, Toshimitsu, et al.
Veröffentlicht: (2024)
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
von: Chowdhury, Mohammed Nowaz Rabbani, et al.
Veröffentlicht: (2024)
von: Chowdhury, Mohammed Nowaz Rabbani, et al.
Veröffentlicht: (2024)
ElasticZO: A Memory-Efficient On-Device Learning with Combined Zeroth- and First-Order Optimization
von: Sugiura, Keisuke, et al.
Veröffentlicht: (2025)
von: Sugiura, Keisuke, et al.
Veröffentlicht: (2025)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
von: Bu, Dake, et al.
Veröffentlicht: (2025)
von: Bu, Dake, et al.
Veröffentlicht: (2025)
Test time training enhances in-context learning of nonlinear functions
von: Kuwataka, Kento, et al.
Veröffentlicht: (2025)
von: Kuwataka, Kento, et al.
Veröffentlicht: (2025)
Deep Two-Way Matrix Reordering for Relational Data Analysis
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
von: Takakura, Shokichi, et al.
Veröffentlicht: (2024)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2024)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023)
Mixture of Experts Meets Prompt-Based Continual Learning
von: Le, Minh, et al.
Veröffentlicht: (2024)
von: Le, Minh, et al.
Veröffentlicht: (2024)
Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations
von: Oko, Kazusato, et al.
Veröffentlicht: (2024)
von: Oko, Kazusato, et al.
Veröffentlicht: (2024)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
von: Park, Sejik
Veröffentlicht: (2024)
von: Park, Sejik
Veröffentlicht: (2024)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
von: Jiang, Jiarui, et al.
Veröffentlicht: (2025)
von: Jiang, Jiarui, et al.
Veröffentlicht: (2025)
Accelerating Local LLMs on Resource-Constrained Edge Devices via Distributed Prompt Caching
von: Matsutani, Hiroki, et al.
Veröffentlicht: (2026)
von: Matsutani, Hiroki, et al.
Veröffentlicht: (2026)
Gradient Descent with Provably Tuned Learning-rate Schedules
von: Sharma, Dravyansh
Veröffentlicht: (2025)
von: Sharma, Dravyansh
Veröffentlicht: (2025)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
von: Bu, Dake, et al.
Veröffentlicht: (2025)
von: Bu, Dake, et al.
Veröffentlicht: (2025)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
Learning Provably Improves the Convergence of Gradient Descent
von: Song, Qingyu, et al.
Veröffentlicht: (2025)
von: Song, Qingyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024) -
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026) -
Direct Distributional Optimization for Provable Alignment of Diffusion Models
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025) -
Dataset Distillation Efficiently Encodes Low-Dimensional Representations from Gradient-Based Learning of Non-Linear Tasks
von: Kinoshita, Yuri, et al.
Veröffentlicht: (2026) -
From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)