Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Hongbo, Wu, Qinhang, Lin, Sen, Liang, Yingbin, Shroff, Ness B.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908622292779008
author Li, Hongbo
Wu, Qinhang
Lin, Sen
Liang, Yingbin
Shroff, Ness B.
author_facet Li, Hongbo
Wu, Qinhang
Lin, Sen
Liang, Yingbin
Shroff, Ness B.
contents Mixture-of-Experts (MoE) models improve transformer efficiency but lack a unified theoretical explanation, especially when both feed-forward and attention layers are allowed to specialize. To this end, we study the Mixture-of-Transformers (MoT), a tractable theoretical framework in which each transformer block acts as an expert governed by a continuously trained gating network. This design allows us to isolate and study the core learning dynamics of expert specialization and attention alignment. In particular, we develop a three-stage training algorithm with continuous training of the gating network, and show that each transformer expert specializes in a distinct class of tasks and that the gating network accurately routes data samples to the correct expert. Our analysis shows how expert specialization reduces gradient conflicts and makes each subtask strongly convex. We prove that the training drives the expected prediction loss to near zero in $O(\log(ε^{-1}))$ iteration steps, significantly improving over the $O(ε^{-1})$ rate for a single transformer. We further validate our theoretical findings through extensive real-data experiments, demonstrating the practical effectiveness of MoT. Together, these results offer the first unified theoretical account of transformer-level specialization and learning dynamics, providing practical guidance for designing efficient large-scale models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
Li, Hongbo
Wu, Qinhang
Lin, Sen
Liang, Yingbin
Shroff, Ness B.
Machine Learning
Mixture-of-Experts (MoE) models improve transformer efficiency but lack a unified theoretical explanation, especially when both feed-forward and attention layers are allowed to specialize. To this end, we study the Mixture-of-Transformers (MoT), a tractable theoretical framework in which each transformer block acts as an expert governed by a continuously trained gating network. This design allows us to isolate and study the core learning dynamics of expert specialization and attention alignment. In particular, we develop a three-stage training algorithm with continuous training of the gating network, and show that each transformer expert specializes in a distinct class of tasks and that the gating network accurately routes data samples to the correct expert. Our analysis shows how expert specialization reduces gradient conflicts and makes each subtask strongly convex. We prove that the training drives the expected prediction loss to near zero in $O(\log(ε^{-1}))$ iteration steps, significantly improving over the $O(ε^{-1})$ rate for a single transformer. We further validate our theoretical findings through extensive real-data experiments, demonstrating the practical effectiveness of MoT. Together, these results offer the first unified theoretical account of transformer-level specialization and learning dynamics, providing practical guidance for designing efficient large-scale models.
title Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
topic Machine Learning
url https://arxiv.org/abs/2510.27004