Open Multimodal Retrieval-Augmented Factual Image Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tian, Yang, Liu, Fan, Zhang, Jingyuan, Bi, Wei, Hu, Yupeng, Nie, Liqiang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918171476230144
author Tian, Yang
Liu, Fan
Zhang, Jingyuan
Bi, Wei
Hu, Yupeng
Nie, Liqiang
author_facet Tian, Yang
Liu, Fan
Zhang, Jingyuan
Bi, Wei
Hu, Yupeng
Nie, Liqiang
contents Large Multimodal Models (LMMs) have achieved remarkable progress in generating photorealistic and prompt-aligned images, but they often produce outputs that contradict verifiable knowledge, especially when prompts involve fine-grained attributes or time-sensitive events. Conventional retrieval-augmented approaches attempt to address this issue by introducing external information, yet they are fundamentally incapable of grounding generation in accurate and evolving knowledge due to their reliance on static sources and shallow evidence integration. To bridge this gap, we introduce ORIG, an agentic open multimodal retrieval-augmented framework for Factual Image Generation (FIG), a new task that requires both visual realism and factual grounding. ORIG iteratively retrieves and filters multimodal evidence from the web and incrementally integrates the refined knowledge into enriched prompts to guide generation. To support systematic evaluation, we build FIG-Eval, a benchmark spanning ten categories across perceptual, compositional, and temporal dimensions. Experiments demonstrate that ORIG substantially improves factual consistency and overall image quality over strong baselines, highlighting the potential of open multimodal retrieval for factual image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22521
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Open Multimodal Retrieval-Augmented Factual Image Generation
Tian, Yang
Liu, Fan
Zhang, Jingyuan
Bi, Wei
Hu, Yupeng
Nie, Liqiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
Machine Learning
Large Multimodal Models (LMMs) have achieved remarkable progress in generating photorealistic and prompt-aligned images, but they often produce outputs that contradict verifiable knowledge, especially when prompts involve fine-grained attributes or time-sensitive events. Conventional retrieval-augmented approaches attempt to address this issue by introducing external information, yet they are fundamentally incapable of grounding generation in accurate and evolving knowledge due to their reliance on static sources and shallow evidence integration. To bridge this gap, we introduce ORIG, an agentic open multimodal retrieval-augmented framework for Factual Image Generation (FIG), a new task that requires both visual realism and factual grounding. ORIG iteratively retrieves and filters multimodal evidence from the web and incrementally integrates the refined knowledge into enriched prompts to guide generation. To support systematic evaluation, we build FIG-Eval, a benchmark spanning ten categories across perceptual, compositional, and temporal dimensions. Experiments demonstrate that ORIG substantially improves factual consistency and overall image quality over strong baselines, highlighting the potential of open multimodal retrieval for factual image generation.
title Open Multimodal Retrieval-Augmented Factual Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2510.22521