VQ-VA World: Towards High-Quality Visual Question-Visual Answering

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gou, Chenhui, Chen, Zilong, Wang, Zeyu, Li, Feng, Zhu, Deyao, Duan, Zicheng, Li, Kunchang, Deng, Chaorui, Yuan, Hongyi, Fan, Haoqi, Xie, Cihang, Cai, Jianfei, Rezatofighi, Hamid
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914170934394880
author Gou, Chenhui
Chen, Zilong
Wang, Zeyu
Li, Feng
Zhu, Deyao
Duan, Zicheng
Li, Kunchang
Deng, Chaorui
Yuan, Hongyi
Fan, Haoqi
Xie, Cihang
Cai, Jianfei
Rezatofighi, Hamid
author_facet Gou, Chenhui
Chen, Zilong
Wang, Zeyu
Li, Feng
Zhu, Deyao
Duan, Zicheng
Li, Kunchang
Deng, Chaorui
Yuan, Hongyi
Fan, Haoqi
Xie, Cihang
Cai, Jianfei
Rezatofighi, Hamid
contents This paper studies Visual Question-Visual Answering (VQ-VA): generating an image, rather than text, in response to a visual question -- an ability that has recently emerged in proprietary systems such as NanoBanana and GPT-Image. To also bring this capability to open-source models, we introduce VQ-VA World, a data-centric framework built around an agentic pipeline for large-scale, targeted data construction. Leveraging web-scale deployment, this pipeline crawls a massive amount of ~1.8M high-quality, interleaved image-text samples for model training. For evaluation, we further release IntelligentBench, a human-curated benchmark that systematically assesses VQ-VA along the aspects of world knowledge, design knowledge, and reasoning. Training with VQ-VA World data yields strong empirical gains: it helps LightFusion attain 53.06 on IntelligentBench, substantially surpassing the best prior open-source baselines (i.e., 7.78 from vanilla LightFusion; 1.94 from UniWorld-V1), and significantly narrowing the gap toward leading proprietary systems (e.g., 81.67 from NanoBanana; 82.64 from GPT-Image). By releasing the full suite of model weights, datasets, and pipelines, we hope to stimulate future research on VQ-VA.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20573
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VQ-VA World: Towards High-Quality Visual Question-Visual Answering
Gou, Chenhui
Chen, Zilong
Wang, Zeyu
Li, Feng
Zhu, Deyao
Duan, Zicheng
Li, Kunchang
Deng, Chaorui
Yuan, Hongyi
Fan, Haoqi
Xie, Cihang
Cai, Jianfei
Rezatofighi, Hamid
Computer Vision and Pattern Recognition
This paper studies Visual Question-Visual Answering (VQ-VA): generating an image, rather than text, in response to a visual question -- an ability that has recently emerged in proprietary systems such as NanoBanana and GPT-Image. To also bring this capability to open-source models, we introduce VQ-VA World, a data-centric framework built around an agentic pipeline for large-scale, targeted data construction. Leveraging web-scale deployment, this pipeline crawls a massive amount of ~1.8M high-quality, interleaved image-text samples for model training. For evaluation, we further release IntelligentBench, a human-curated benchmark that systematically assesses VQ-VA along the aspects of world knowledge, design knowledge, and reasoning. Training with VQ-VA World data yields strong empirical gains: it helps LightFusion attain 53.06 on IntelligentBench, substantially surpassing the best prior open-source baselines (i.e., 7.78 from vanilla LightFusion; 1.94 from UniWorld-V1), and significantly narrowing the gap toward leading proprietary systems (e.g., 81.67 from NanoBanana; 82.64 from GPT-Image). By releasing the full suite of model weights, datasets, and pipelines, we hope to stimulate future research on VQ-VA.
title VQ-VA World: Towards High-Quality Visual Question-Visual Answering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.20573