Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Fangda, Xie, Zhifei, Hu, Yuxin, Yin, Yihang, Huang, Shurui, Dong, Shikai, Bao, Jianzhu, Yan, Shuicheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908977556619264
author Ye, Fangda
Xie, Zhifei
Hu, Yuxin
Yin, Yihang
Huang, Shurui
Dong, Shikai
Bao, Jianzhu
Yan, Shuicheng
author_facet Ye, Fangda
Xie, Zhifei
Hu, Yuxin
Yin, Yihang
Huang, Shurui
Dong, Shikai
Bao, Jianzhu
Yan, Shuicheng
contents Recent agentic search frameworks enable deep research via iterative planning and retrieval, reducing hallucinations and enhancing factual grounding. However, they remain text-centric, overlooking the multimodal evidence that characterizes real-world expert reports. We introduce a pressing task: multimodal long-form generation. Accordingly, we propose Deep-Reporter, a unified agentic framework for grounded multimodal long-form generation. It orchestrates: (i) Agentic Multimodal Search and Filtering to retrieve and filter textual passages and information-dense visuals; (ii) Checklist-Guided Incremental Synthesis to ensure coherent image-text integration and optimal citation placement; and (iii) Recurrent Context Management to balance long-range coherence with local fluency. We develop a rigorous curation pipeline producing 8K high-quality agentic traces for model optimization. We further introduce M2LongBench, a comprehensive testbed comprising 247 research tasks across 9 domains and a stable multimodal sandbox. Extensive experiments demonstrate that long-form multimodal generation is a challenging task, especially in multimodal selection and integration, and effective post-training can bridge the gap.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10741
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
Ye, Fangda
Xie, Zhifei
Hu, Yuxin
Yin, Yihang
Huang, Shurui
Dong, Shikai
Bao, Jianzhu
Yan, Shuicheng
Computation and Language
Artificial Intelligence
Information Retrieval
Recent agentic search frameworks enable deep research via iterative planning and retrieval, reducing hallucinations and enhancing factual grounding. However, they remain text-centric, overlooking the multimodal evidence that characterizes real-world expert reports. We introduce a pressing task: multimodal long-form generation. Accordingly, we propose Deep-Reporter, a unified agentic framework for grounded multimodal long-form generation. It orchestrates: (i) Agentic Multimodal Search and Filtering to retrieve and filter textual passages and information-dense visuals; (ii) Checklist-Guided Incremental Synthesis to ensure coherent image-text integration and optimal citation placement; and (iii) Recurrent Context Management to balance long-range coherence with local fluency. We develop a rigorous curation pipeline producing 8K high-quality agentic traces for model optimization. We further introduce M2LongBench, a comprehensive testbed comprising 247 research tasks across 9 domains and a stable multimodal sandbox. Extensive experiments demonstrate that long-form multimodal generation is a challenging task, especially in multimodal selection and integration, and effective post-training can bridge the gap.
title Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2604.10741