Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908977556619264 |
|---|---|
| author | Ye, Fangda Xie, Zhifei Hu, Yuxin Yin, Yihang Huang, Shurui Dong, Shikai Bao, Jianzhu Yan, Shuicheng |
| author_facet | Ye, Fangda Xie, Zhifei Hu, Yuxin Yin, Yihang Huang, Shurui Dong, Shikai Bao, Jianzhu Yan, Shuicheng |
| contents | Recent agentic search frameworks enable deep research via iterative planning and retrieval, reducing hallucinations and enhancing factual grounding. However, they remain text-centric, overlooking the multimodal evidence that characterizes real-world expert reports. We introduce a pressing task: multimodal long-form generation. Accordingly, we propose Deep-Reporter, a unified agentic framework for grounded multimodal long-form generation. It orchestrates: (i) Agentic Multimodal Search and Filtering to retrieve and filter textual passages and information-dense visuals; (ii) Checklist-Guided Incremental Synthesis to ensure coherent image-text integration and optimal citation placement; and (iii) Recurrent Context Management to balance long-range coherence with local fluency. We develop a rigorous curation pipeline producing 8K high-quality agentic traces for model optimization. We further introduce M2LongBench, a comprehensive testbed comprising 247 research tasks across 9 domains and a stable multimodal sandbox. Extensive experiments demonstrate that long-form multimodal generation is a challenging task, especially in multimodal selection and integration, and effective post-training can bridge the gap. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_10741 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation Ye, Fangda Xie, Zhifei Hu, Yuxin Yin, Yihang Huang, Shurui Dong, Shikai Bao, Jianzhu Yan, Shuicheng Computation and Language Artificial Intelligence Information Retrieval Recent agentic search frameworks enable deep research via iterative planning and retrieval, reducing hallucinations and enhancing factual grounding. However, they remain text-centric, overlooking the multimodal evidence that characterizes real-world expert reports. We introduce a pressing task: multimodal long-form generation. Accordingly, we propose Deep-Reporter, a unified agentic framework for grounded multimodal long-form generation. It orchestrates: (i) Agentic Multimodal Search and Filtering to retrieve and filter textual passages and information-dense visuals; (ii) Checklist-Guided Incremental Synthesis to ensure coherent image-text integration and optimal citation placement; and (iii) Recurrent Context Management to balance long-range coherence with local fluency. We develop a rigorous curation pipeline producing 8K high-quality agentic traces for model optimization. We further introduce M2LongBench, a comprehensive testbed comprising 247 research tasks across 9 domains and a stable multimodal sandbox. Extensive experiments demonstrate that long-form multimodal generation is a challenging task, especially in multimodal selection and integration, and effective post-training can bridge the gap. |
| title | Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation |
| topic | Computation and Language Artificial Intelligence Information Retrieval |
| url | https://arxiv.org/abs/2604.10741 |