All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rahman, Tanzila, Liao, Renjie, Sigal, Leonid
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914470993854464
author Rahman, Tanzila
Liao, Renjie
Sigal, Leonid
author_facet Rahman, Tanzila
Liao, Renjie
Sigal, Leonid
contents Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating multimodal video data in real-world is costly, slow, and inherently limited in diversity and coverage. To address this challenge, we propose a unified synthetic data generation pipeline capable of automatically producing unlimited multimodal video data with rich and diverse supervision. Our framework supports multiple task formats within a single pipeline, enabling scalable and consistent data creation across tasks. To further enhance reasoning ability, we introduce a VQA-based fine-tuning strategy that trains models to answer structured questions about visual content rather than relying solely on captions or simple instructions. This formulation encourages deeper visual grounding and reasoning. We evaluate our approach in three challenging tasks: video object counting, video-based visual question answering, and video object segmentation. Experimental results demonstrate that models trained predominantly on synthetic data generalize effectively to real-world datasets, often outperforming traditionally trained counterparts. Our findings highlight the potential of unified synthetic data pipelines as a scalable alternative to expensive real-world annotation for multimodal video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12335
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
Rahman, Tanzila
Liao, Renjie
Sigal, Leonid
Computer Vision and Pattern Recognition
Machine Learning
Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating multimodal video data in real-world is costly, slow, and inherently limited in diversity and coverage. To address this challenge, we propose a unified synthetic data generation pipeline capable of automatically producing unlimited multimodal video data with rich and diverse supervision. Our framework supports multiple task formats within a single pipeline, enabling scalable and consistent data creation across tasks. To further enhance reasoning ability, we introduce a VQA-based fine-tuning strategy that trains models to answer structured questions about visual content rather than relying solely on captions or simple instructions. This formulation encourages deeper visual grounding and reasoning. We evaluate our approach in three challenging tasks: video object counting, video-based visual question answering, and video object segmentation. Experimental results demonstrate that models trained predominantly on synthetic data generalize effectively to real-world datasets, often outperforming traditionally trained counterparts. Our findings highlight the potential of unified synthetic data pipelines as a scalable alternative to expensive real-world annotation for multimodal video understanding.
title All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2604.12335