Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Honglin, Qin, Chonghan, Liu, Zheng, Pei, Qizhi, Li, Yu, Zhong, Zhanping, Gao, Xin, Wang, Yanfeng, He, Conghui, Wu, Lijun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912845113851904
author Lin, Honglin
Qin, Chonghan
Liu, Zheng
Pei, Qizhi
Li, Yu
Zhong, Zhanping
Gao, Xin
Wang, Yanfeng
He, Conghui
Wu, Lijun
author_facet Lin, Honglin
Qin, Chonghan
Liu, Zheng
Pei, Qizhi
Li, Yu
Zhong, Zhanping
Gao, Xin
Wang, Yanfeng
He, Conghui
Wu, Lijun
contents While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained by the difficulty of synthesizing scientifically rigorous images. Existing Text-to-Image (T2I) models often produce outputs that are visually plausible yet scientifically incorrect, resulting in a persistent visual-logic divergence that limits their value for downstream reasoning. Motivated by recent advances in next-generation T2I models, we conduct a systematic study of scientific image synthesis across generation paradigms, evaluation, and downstream use. We analyze both direct pixel-based generation and programmatic synthesis, and propose ImgCoder, a logic-driven framework that follows an explicit "understand - plan - code" workflow to improve structural precision. To rigorously assess scientific correctness, we introduce SciGenBench, which evaluates generated images based on information utility and logical validity. Our evaluation reveals systematic failure modes in pixel-based models and highlights a fundamental expressiveness-precision trade-off. Finally, we show that fine-tuning Large Multimodal Models (LMMs) on rigorously verified synthetic scientific images yields consistent reasoning gains, with potential scaling trends analogous to the text domain, validating high-fidelity scientific synthesis as a viable path to unlocking massive multimodal reasoning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17027
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
Lin, Honglin
Qin, Chonghan
Liu, Zheng
Pei, Qizhi
Li, Yu
Zhong, Zhanping
Gao, Xin
Wang, Yanfeng
He, Conghui
Wu, Lijun
Computer Vision and Pattern Recognition
Artificial Intelligence
While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained by the difficulty of synthesizing scientifically rigorous images. Existing Text-to-Image (T2I) models often produce outputs that are visually plausible yet scientifically incorrect, resulting in a persistent visual-logic divergence that limits their value for downstream reasoning. Motivated by recent advances in next-generation T2I models, we conduct a systematic study of scientific image synthesis across generation paradigms, evaluation, and downstream use. We analyze both direct pixel-based generation and programmatic synthesis, and propose ImgCoder, a logic-driven framework that follows an explicit "understand - plan - code" workflow to improve structural precision. To rigorously assess scientific correctness, we introduce SciGenBench, which evaluates generated images based on information utility and logical validity. Our evaluation reveals systematic failure modes in pixel-based models and highlights a fundamental expressiveness-precision trade-off. Finally, we show that fine-tuning Large Multimodal Models (LMMs) on rigorously verified synthetic scientific images yields consistent reasoning gains, with potential scaling trends analogous to the text domain, validating high-fidelity scientific synthesis as a viable path to unlocking massive multimodal reasoning capabilities.
title Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.17027