UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909860409376768 |
|---|---|
| author | Wang, Fu-Yun Zhang, Han Gharbi, Michael Li, Hongsheng Park, Taesung |
| author_facet | Wang, Fu-Yun Zhang, Han Gharbi, Michael Li, Hongsheng Park, Taesung |
| contents | We present UniRL-Zero, a unified reinforcement learning (RL) framework that boosts, multimodal language model understanding and reasoning, diffusion model multimedia generation, and their beneficial interaction capabilities within a unified model. Our work defines six scenarios for unified model reinforcement learning, providing systematic baselines for reinforcement learning of unified understanding and generation model. Our code is available at https://github.com/G-U-N/UniRL. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_17937 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts Wang, Fu-Yun Zhang, Han Gharbi, Michael Li, Hongsheng Park, Taesung Machine Learning Artificial Intelligence We present UniRL-Zero, a unified reinforcement learning (RL) framework that boosts, multimodal language model understanding and reasoning, diffusion model multimedia generation, and their beneficial interaction capabilities within a unified model. Our work defines six scenarios for unified model reinforcement learning, providing systematic baselines for reinforcement learning of unified understanding and generation model. Our code is available at https://github.com/G-U-N/UniRL. |
| title | UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2510.17937 |