UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Fu-Yun, Zhang, Han, Gharbi, Michael, Li, Hongsheng, Park, Taesung
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909860409376768
author Wang, Fu-Yun
Zhang, Han
Gharbi, Michael
Li, Hongsheng
Park, Taesung
author_facet Wang, Fu-Yun
Zhang, Han
Gharbi, Michael
Li, Hongsheng
Park, Taesung
contents We present UniRL-Zero, a unified reinforcement learning (RL) framework that boosts, multimodal language model understanding and reasoning, diffusion model multimedia generation, and their beneficial interaction capabilities within a unified model. Our work defines six scenarios for unified model reinforcement learning, providing systematic baselines for reinforcement learning of unified understanding and generation model. Our code is available at https://github.com/G-U-N/UniRL.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17937
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
Wang, Fu-Yun
Zhang, Han
Gharbi, Michael
Li, Hongsheng
Park, Taesung
Machine Learning
Artificial Intelligence
We present UniRL-Zero, a unified reinforcement learning (RL) framework that boosts, multimodal language model understanding and reasoning, diffusion model multimedia generation, and their beneficial interaction capabilities within a unified model. Our work defines six scenarios for unified model reinforcement learning, providing systematic baselines for reinforcement learning of unified understanding and generation model. Our code is available at https://github.com/G-U-N/UniRL.
title UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.17937