ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Chaoyu, Kulkarni, Yogesh, Fazli, Pooyan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914514753028096
author Li, Chaoyu
Kulkarni, Yogesh
Fazli, Pooyan
author_facet Li, Chaoyu
Kulkarni, Yogesh
Fazli, Pooyan
contents The computational cost of training multimodal large language models (MLLMs) grows rapidly with the number of processed tokens. Existing efficiency methods mainly target inference via token reduction or merging, offering limited benefits during training. We introduce ReGATE (Reference-Guided Adaptive Token Elision), an adaptive token pruning method for accelerating MLLM training. ReGATE adopts a teacher-student framework, in which a frozen teacher LLM provides per-token guidance losses that are fused with an exponential moving average of the student's difficulty estimates. This adaptive scoring mechanism dynamically selects informative tokens while skipping redundant ones in the forward pass, substantially reducing computation without altering the model architecture. Across three representative MLLMs, ReGATE matches the peak accuracy of standard training on MVBench up to 2$\times$ faster, using only 38% of the tokens. With extended training, it even surpasses the baseline across multiple multimodal benchmarks, cutting total token usage by over 41%.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21420
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
Li, Chaoyu
Kulkarni, Yogesh
Fazli, Pooyan
Computer Vision and Pattern Recognition
Computation and Language
The computational cost of training multimodal large language models (MLLMs) grows rapidly with the number of processed tokens. Existing efficiency methods mainly target inference via token reduction or merging, offering limited benefits during training. We introduce ReGATE (Reference-Guided Adaptive Token Elision), an adaptive token pruning method for accelerating MLLM training. ReGATE adopts a teacher-student framework, in which a frozen teacher LLM provides per-token guidance losses that are fused with an exponential moving average of the student's difficulty estimates. This adaptive scoring mechanism dynamically selects informative tokens while skipping redundant ones in the forward pass, substantially reducing computation without altering the model architecture. Across three representative MLLMs, ReGATE matches the peak accuracy of standard training on MVBench up to 2$\times$ faster, using only 38% of the tokens. With extended training, it even surpasses the baseline across multiple multimodal benchmarks, cutting total token usage by over 41%.
title ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2507.21420