EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Thawakar, Omkar, Venkatraman, Shravan, Thawkar, Ritesh, Shaker, Abdelrahman, Cholakkal, Hisham, Anwer, Rao Muhammad, Khan, Salman, Khan, Fahad
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917337104384000
author Thawakar, Omkar
Venkatraman, Shravan
Thawkar, Ritesh
Shaker, Abdelrahman
Cholakkal, Hisham
Anwer, Rao Muhammad
Khan, Salman
Khan, Fahad
author_facet Thawakar, Omkar
Venkatraman, Shravan
Thawkar, Ritesh
Shaker, Abdelrahman
Cholakkal, Hisham
Anwer, Rao Muhammad
Khan, Salman
Khan, Fahad
contents Recent advances in large multimodal models (LMMs) have enabled impressive reasoning and perception abilities, yet most existing training pipelines still depend on human-curated data or externally verified reward models, limiting their autonomy and scalability. In this work, we strive to improve LMM reasoning capabilities in a purely unsupervised fashion (without any annotated data or reward distillation). To this end, we propose a self-evolving framework, named EvoLMM, that instantiates two cooperative agents from a single backbone model: a Proposer, which generates diverse, image-grounded questions, and a Solver, which solves them through internal consistency, where learning proceeds through a continuous self-rewarding process. This dynamic feedback encourages both the generation of informative queries and the refinement of structured reasoning without relying on ground-truth or human judgments. When using the popular Qwen2.5-VL as the base model, our EvoLMM yields consistent gains upto $\sim$3\% on multimodal math-reasoning benchmarks, including ChartQA, MathVista, and MathVision, using only raw training images. We hope our simple yet effective approach will serve as a solid baseline easing future research in self-improving LMMs in a fully-unsupervised fashion. Our code and models are available at https://github.com/mbzuai-oryx/EvoLMM.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16672
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
Thawakar, Omkar
Venkatraman, Shravan
Thawkar, Ritesh
Shaker, Abdelrahman
Cholakkal, Hisham
Anwer, Rao Muhammad
Khan, Salman
Khan, Fahad
Computer Vision and Pattern Recognition
Recent advances in large multimodal models (LMMs) have enabled impressive reasoning and perception abilities, yet most existing training pipelines still depend on human-curated data or externally verified reward models, limiting their autonomy and scalability. In this work, we strive to improve LMM reasoning capabilities in a purely unsupervised fashion (without any annotated data or reward distillation). To this end, we propose a self-evolving framework, named EvoLMM, that instantiates two cooperative agents from a single backbone model: a Proposer, which generates diverse, image-grounded questions, and a Solver, which solves them through internal consistency, where learning proceeds through a continuous self-rewarding process. This dynamic feedback encourages both the generation of informative queries and the refinement of structured reasoning without relying on ground-truth or human judgments. When using the popular Qwen2.5-VL as the base model, our EvoLMM yields consistent gains upto $\sim$3\% on multimodal math-reasoning benchmarks, including ChartQA, MathVista, and MathVision, using only raw training images. We hope our simple yet effective approach will serve as a solid baseline easing future research in self-improving LMMs in a fully-unsupervised fashion. Our code and models are available at https://github.com/mbzuai-oryx/EvoLMM.
title EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.16672