Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Kunyu, Ma, Yue, Zhang, Xinhua, Liu, Boshi, Yuluo, Yikuang, Zhang, Yinhan, Liu, Runtao, Liu, Hongyu, Qin, Zhiyuan, Mo, Shanhui, Chen, Qifeng, Wang, Zeyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915433871835136
author Feng, Kunyu
Ma, Yue
Zhang, Xinhua
Liu, Boshi
Yuluo, Yikuang
Zhang, Yinhan
Liu, Runtao
Liu, Hongyu
Qin, Zhiyuan
Mo, Shanhui
Chen, Qifeng
Wang, Zeyu
author_facet Feng, Kunyu
Ma, Yue
Zhang, Xinhua
Liu, Boshi
Yuluo, Yikuang
Zhang, Yinhan
Liu, Runtao
Liu, Hongyu
Qin, Zhiyuan
Mo, Shanhui
Chen, Qifeng
Wang, Zeyu
contents With the growing demands of AI-generated content (AIGC), the need for high-quality, diverse, and scalable data has become increasingly crucial. However, collecting large-scale real-world data remains costly and time-consuming, hindering the development of downstream applications. While some works attempt to collect task-specific data via a rendering process, most approaches still rely on manual scene construction, limiting their scalability and accuracy. To address these challenges, we propose Follow-Your-Instruction, a Multimodal Large Language Model (MLLM)-driven framework for automatically synthesizing high-quality 2D, 3D, and 4D data. Our \textbf{Follow-Your-Instruction} first collects assets and their associated descriptions through multimodal inputs using the MLLM-Collector. Then it constructs 3D layouts, and leverages Vision-Language Models (VLMs) for semantic refinement through multi-view scenes with the MLLM-Generator and MLLM-Optimizer, respectively. Finally, it uses MLLM-Planner to generate temporally coherent future frames. We evaluate the quality of the generated data through comprehensive experiments on the 2D, 3D, and 4D generative tasks. The results show that our synthetic data significantly boosts the performance of existing baseline models, demonstrating Follow-Your-Instruction's potential as a scalable and effective data engine for generative intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05580
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
Feng, Kunyu
Ma, Yue
Zhang, Xinhua
Liu, Boshi
Yuluo, Yikuang
Zhang, Yinhan
Liu, Runtao
Liu, Hongyu
Qin, Zhiyuan
Mo, Shanhui
Chen, Qifeng
Wang, Zeyu
Computer Vision and Pattern Recognition
With the growing demands of AI-generated content (AIGC), the need for high-quality, diverse, and scalable data has become increasingly crucial. However, collecting large-scale real-world data remains costly and time-consuming, hindering the development of downstream applications. While some works attempt to collect task-specific data via a rendering process, most approaches still rely on manual scene construction, limiting their scalability and accuracy. To address these challenges, we propose Follow-Your-Instruction, a Multimodal Large Language Model (MLLM)-driven framework for automatically synthesizing high-quality 2D, 3D, and 4D data. Our \textbf{Follow-Your-Instruction} first collects assets and their associated descriptions through multimodal inputs using the MLLM-Collector. Then it constructs 3D layouts, and leverages Vision-Language Models (VLMs) for semantic refinement through multi-view scenes with the MLLM-Generator and MLLM-Optimizer, respectively. Finally, it uses MLLM-Planner to generate temporally coherent future frames. We evaluate the quality of the generated data through comprehensive experiments on the 2D, 3D, and 4D generative tasks. The results show that our synthetic data significantly boosts the performance of existing baseline models, demonstrating Follow-Your-Instruction's potential as a scalable and effective data engine for generative intelligence.
title Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.05580