Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tong, Jingqi, Tang, Jixin, Li, Hangcheng, Mou, Yurong, Zhang, Ming, Zhao, Jun, Wen, Yanbo, Song, Fan, Zhan, Jiahao, Lu, Yuyang, Tao, Chaoran, Guo, Zhiyuan, Yu, Jizhou, Cheng, Tianhao, Xi, Zhiheng, Jiang, Changhao, Yin, Zhangyue, Zheng, Yining, Ge, Weifeng, Chen, Guanhua, Gui, Tao, Qiu, Xipeng, Zhang, Qi, Huang, Xuanjing
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908705159643136
author Tong, Jingqi
Tang, Jixin
Li, Hangcheng
Mou, Yurong
Zhang, Ming
Zhao, Jun
Wen, Yanbo
Song, Fan
Zhan, Jiahao
Lu, Yuyang
Tao, Chaoran
Guo, Zhiyuan
Yu, Jizhou
Cheng, Tianhao
Xi, Zhiheng
Jiang, Changhao
Yin, Zhangyue
Zheng, Yining
Ge, Weifeng
Chen, Guanhua
Gui, Tao
Qiu, Xipeng
Zhang, Qi
Huang, Xuanjing
author_facet Tong, Jingqi
Tang, Jixin
Li, Hangcheng
Mou, Yurong
Zhang, Ming
Zhao, Jun
Wen, Yanbo
Song, Fan
Zhan, Jiahao
Lu, Yuyang
Tao, Chaoran
Guo, Zhiyuan
Yu, Jizhou
Cheng, Tianhao
Xi, Zhiheng
Jiang, Changhao
Yin, Zhangyue
Zheng, Yining
Ge, Weifeng
Chen, Guanhua
Gui, Tao
Qiu, Xipeng
Zhang, Qi
Huang, Xuanjing
contents Vision-language reinforcement learning (RL) has primarily focused on narrow domains (e.g. geometry or chart reasoning). This leaves broader training scenarios and resources underexplored, limiting the exploration and learning of Vision Language Models (VLMs) through RL. We find video games inherently provide rich visual elements and mechanics that are easy to verify. To fully use the multimodal and verifiable reward in video games, we propose Game-RL, constructing diverse game tasks for RL training to boost VLMs general reasoning ability. To obtain training data, we propose Code2Logic, a novel approach that adapts game code to synthesize game reasoning task data, thus obtaining the GameQA dataset of 30 games and 158 tasks with controllable difficulty gradation. Unexpectedly, RL training solely on GameQA enables multiple VLMs to achieve performance improvements across 7 diverse vision-language benchmarks, demonstrating the value of Game-RL for enhancing VLMs' general reasoning. Furthermore, this suggests that video games may serve as valuable scenarios and resources to boost general reasoning abilities. Our code, dataset and models are available at the GitHub repository.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13886
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
Tong, Jingqi
Tang, Jixin
Li, Hangcheng
Mou, Yurong
Zhang, Ming
Zhao, Jun
Wen, Yanbo
Song, Fan
Zhan, Jiahao
Lu, Yuyang
Tao, Chaoran
Guo, Zhiyuan
Yu, Jizhou
Cheng, Tianhao
Xi, Zhiheng
Jiang, Changhao
Yin, Zhangyue
Zheng, Yining
Ge, Weifeng
Chen, Guanhua
Gui, Tao
Qiu, Xipeng
Zhang, Qi
Huang, Xuanjing
Computation and Language
I.2.7; I.2.10
Vision-language reinforcement learning (RL) has primarily focused on narrow domains (e.g. geometry or chart reasoning). This leaves broader training scenarios and resources underexplored, limiting the exploration and learning of Vision Language Models (VLMs) through RL. We find video games inherently provide rich visual elements and mechanics that are easy to verify. To fully use the multimodal and verifiable reward in video games, we propose Game-RL, constructing diverse game tasks for RL training to boost VLMs general reasoning ability. To obtain training data, we propose Code2Logic, a novel approach that adapts game code to synthesize game reasoning task data, thus obtaining the GameQA dataset of 30 games and 158 tasks with controllable difficulty gradation. Unexpectedly, RL training solely on GameQA enables multiple VLMs to achieve performance improvements across 7 diverse vision-language benchmarks, demonstrating the value of Game-RL for enhancing VLMs' general reasoning. Furthermore, this suggests that video games may serve as valuable scenarios and resources to boost general reasoning abilities. Our code, dataset and models are available at the GitHub repository.
title Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
topic Computation and Language
I.2.7; I.2.10
url https://arxiv.org/abs/2505.13886