Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongbo, Wu, Qinhang, Lin, Sen, Liang, Yingbin, Shroff, Ness B. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlocking the Power of Rehearsal in Continual Learning: A Theoretical Perspective
by: Deng, Junze, et al.
Published: (2025)
by: Deng, Junze, et al.
Published: (2025)
Constraint-Rectified Training for Efficient Chain-of-Thought
by: Wu, Qinhang, et al.
Published: (2026)
by: Wu, Qinhang, et al.
Published: (2026)
Theory on Mixture-of-Experts in Continual Learning
by: Li, Hongbo, et al.
Published: (2024)
by: Li, Hongbo, et al.
Published: (2024)
From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models
by: Liang, Yuchen, et al.
Published: (2026)
by: Liang, Yuchen, et al.
Published: (2026)
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
by: Shi, Ming, et al.
Published: (2023)
by: Shi, Ming, et al.
Published: (2023)
Can We Theoretically Quantify the Impacts of Local Updates on the Generalization Performance of Federated Learning?
by: Ju, Peizhong, et al.
Published: (2024)
by: Ju, Peizhong, et al.
Published: (2024)
Monitoring State Transitions in Markovian Systems with Sampling Cost
by: Saurav, Kumar, et al.
Published: (2025)
by: Saurav, Kumar, et al.
Published: (2025)
Regret Bounds for Reinforcement Learning from Multi-Source Imperfect Preferences
by: Shi, Ming, et al.
Published: (2026)
by: Shi, Ming, et al.
Published: (2026)
Theory on Score-Mismatched Diffusion Models and Zero-Shot Conditional Samplers
by: Liang, Yuchen, et al.
Published: (2024)
by: Liang, Yuchen, et al.
Published: (2024)
Sharp Convergence Rates for Masked Diffusion Models
by: Liang, Yuchen, et al.
Published: (2026)
by: Liang, Yuchen, et al.
Published: (2026)
Discrete Diffusion Models: Novel Analysis and New Sampler Guarantees
by: Liang, Yuchen, et al.
Published: (2025)
by: Liang, Yuchen, et al.
Published: (2025)
Broadening Target Distributions for Accelerated Diffusion Models via a Novel Analysis Approach
by: Liang, Yuchen, et al.
Published: (2024)
by: Liang, Yuchen, et al.
Published: (2024)
Provable In-Context Learning of Nonlinear Regression with Transformers
by: Li, Hongbo, et al.
Published: (2025)
by: Li, Hongbo, et al.
Published: (2025)
Absorb and Converge: Provable Convergence Guarantee for Absorbing Discrete Diffusion Models
by: Liang, Yuchen, et al.
Published: (2025)
by: Liang, Yuchen, et al.
Published: (2025)
Learnable Chernoff Baselines for Inference-Time Alignment
by: Madhow, Sunil, et al.
Published: (2026)
by: Madhow, Sunil, et al.
Published: (2026)
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
by: Huang, Ruiquan, et al.
Published: (2025)
by: Huang, Ruiquan, et al.
Published: (2025)
How to Find the Exact Pareto Front for Multi-Objective MDPs?
by: Li, Yining, et al.
Published: (2024)
by: Li, Yining, et al.
Published: (2024)
A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
by: Huang, Yu, et al.
Published: (2024)
by: Huang, Yu, et al.
Published: (2024)
BeST -- A Novel Source Selection Metric for Transfer Learning
by: Soni, Ashutosh, et al.
Published: (2025)
by: Soni, Ashutosh, et al.
Published: (2025)
Provably Efficient Multi-Objective Bandit Algorithms under Preference-Centric Customization
by: Cao, Linfeng, et al.
Published: (2025)
by: Cao, Linfeng, et al.
Published: (2025)
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation
by: Pan, Pei-Chi, et al.
Published: (2026)
by: Pan, Pei-Chi, et al.
Published: (2026)
Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual
by: Li, Yining, et al.
Published: (2026)
by: Li, Yining, et al.
Published: (2026)
Algorithm Design for Online Meta-Learning with Task Boundary Detection
by: Sow, Daouda, et al.
Published: (2023)
by: Sow, Daouda, et al.
Published: (2023)
Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration
by: Roknilamouki, Amirhossein, et al.
Published: (2026)
by: Roknilamouki, Amirhossein, et al.
Published: (2026)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
by: Luo, Ziyue, et al.
Published: (2025)
by: Luo, Ziyue, et al.
Published: (2025)
An LP-based Sampling Policy for Multi-Armed Bandits with Side-Observations and Stochastic Availability
by: Soni, Ashutosh, et al.
Published: (2026)
by: Soni, Ashutosh, et al.
Published: (2026)
Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
by: Zhao, Jinze, et al.
Published: (2024)
by: Zhao, Jinze, et al.
Published: (2024)
AI‐EDGE: An NSF AI institute for future edge networks and distributed intelligence
by: Peizhong Ju, et al.
Published: (2024)
by: Peizhong Ju, et al.
Published: (2024)
FIRM: Federated In-client Regularized Multi-objective Alignment for Large Language Models
by: Nourzad, Fatemeh, et al.
Published: (2025)
by: Nourzad, Fatemeh, et al.
Published: (2025)
Non-asymptotic Convergence of Training Transformers for Next-token Prediction
by: Huang, Ruiquan, et al.
Published: (2024)
by: Huang, Ruiquan, et al.
Published: (2024)
Online Learning for Optimizing AoI-Energy Tradeoff under Unknown Channel Statistics
by: Abd-Elmagid, Mohamed A., et al.
Published: (2025)
by: Abd-Elmagid, Mohamed A., et al.
Published: (2025)
GPT-4o as the Gold Standard: A Scalable and General Purpose Approach to Filter Language Model Pretraining Data
by: Zhang, Jifan, et al.
Published: (2024)
by: Zhang, Jifan, et al.
Published: (2024)
To Theoretically Understand Transformer-Based In-Context Learning for Optimizing CSMA
by: Hao, Shugang, et al.
Published: (2025)
by: Hao, Shugang, et al.
Published: (2025)
Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces
by: Roknilamouki, Amirhossein, et al.
Published: (2025)
by: Roknilamouki, Amirhossein, et al.
Published: (2025)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Theoretical Insights for Diffusion Guidance: A Case Study for Gaussian Mixture Models
by: Wu, Yuchen, et al.
Published: (2024)
by: Wu, Yuchen, et al.
Published: (2024)
Robust Offline Reinforcement Learning for Non-Markovian Decision Processes
by: Huang, Ruiquan, et al.
Published: (2024)
by: Huang, Ruiquan, et al.
Published: (2024)
Theory of Mixture-of-Experts for Mobile Edge Computing
by: Li, Hongbo, et al.
Published: (2024)
by: Li, Hongbo, et al.
Published: (2024)
Provably Efficient UCB-type Algorithms For Learning Predictive State Representations
by: Huang, Ruiquan, et al.
Published: (2023)
by: Huang, Ruiquan, et al.
Published: (2023)
Beyond Freshness and Semantics: A Coupon-Collector Framework for Effective Status Updates
by: Ahmed, Youssef, et al.
Published: (2026)
by: Ahmed, Youssef, et al.
Published: (2026)
Similar Items
-
Unlocking the Power of Rehearsal in Continual Learning: A Theoretical Perspective
by: Deng, Junze, et al.
Published: (2025) -
Constraint-Rectified Training for Efficient Chain-of-Thought
by: Wu, Qinhang, et al.
Published: (2026) -
Theory on Mixture-of-Experts in Continual Learning
by: Li, Hongbo, et al.
Published: (2024) -
From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models
by: Liang, Yuchen, et al.
Published: (2026) -
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
by: Shi, Ming, et al.
Published: (2023)