MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhao, Haozhe, Cai, Zefan, Si, Shuzheng, Chen, Liang, Gu, Jiuxiang, Xiao, Wen, Zhang, Minjia, Hu, Junjie
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913168907829248
author Zhao, Haozhe
Cai, Zefan
Si, Shuzheng
Chen, Liang
Gu, Jiuxiang
Xiao, Wen
Zhang, Minjia
Hu, Junjie
author_facet Zhao, Haozhe
Cai, Zefan
Si, Shuzheng
Chen, Liang
Gu, Jiuxiang
Xiao, Wen
Zhang, Minjia
Hu, Junjie
contents Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requiring extensive training for complex multimodal image generation. To address these limitations, we propose MENTOR, a novel autoregressive (AR) framework for efficient Multimodal-conditioned Tuning for Autoregressive multimodal image generation. MENTOR combines an AR image generator with a two-stage training paradigm, enabling fine-grained, token-level alignment between multimodal inputs and image outputs without relying on auxiliary adapters or cross-attention modules. The two-stage training consists of: (1) a multimodal alignment stage that establishes robust pixel- and semantic-level alignment, followed by (2) a multimodal instruction tuning stage that balances the integration of multimodal inputs and enhances generation controllability. Despite modest model size, suboptimal base components, and limited training resources, MENTOR achieves strong performance on the DreamBench++ benchmark, outperforming competitive baselines in concept preservation and prompt following. Additionally, our method delivers superior image reconstruction fidelity, broad task adaptability, and improved training efficiency compared to diffusion-based methods. Dataset, code, and models are available at: https://github.com/HaozheZhao/MENTOR
format Preprint
id arxiv_https___arxiv_org_abs_2507_09574
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
Zhao, Haozhe
Cai, Zefan
Si, Shuzheng
Chen, Liang
Gu, Jiuxiang
Xiao, Wen
Zhang, Minjia
Hu, Junjie
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requiring extensive training for complex multimodal image generation. To address these limitations, we propose MENTOR, a novel autoregressive (AR) framework for efficient Multimodal-conditioned Tuning for Autoregressive multimodal image generation. MENTOR combines an AR image generator with a two-stage training paradigm, enabling fine-grained, token-level alignment between multimodal inputs and image outputs without relying on auxiliary adapters or cross-attention modules. The two-stage training consists of: (1) a multimodal alignment stage that establishes robust pixel- and semantic-level alignment, followed by (2) a multimodal instruction tuning stage that balances the integration of multimodal inputs and enhances generation controllability. Despite modest model size, suboptimal base components, and limited training resources, MENTOR achieves strong performance on the DreamBench++ benchmark, outperforming competitive baselines in concept preservation and prompt following. Additionally, our method delivers superior image reconstruction fidelity, broad task adaptability, and improved training efficiency compared to diffusion-based methods. Dataset, code, and models are available at: https://github.com/HaozheZhao/MENTOR
title MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.09574