All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Rahman, Tanzila, Liao, Renjie, Sigal, Leonid |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
by: Fan, Wan-Cyuan, et al.
Published: (2024)
by: Fan, Wan-Cyuan, et al.
Published: (2024)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
by: Bhatt, Gaurav, et al.
Published: (2024)
by: Bhatt, Gaurav, et al.
Published: (2024)
Unified Classification and Rejection: A One-versus-All Framework
by: Cheng, Zhen, et al.
Published: (2023)
by: Cheng, Zhen, et al.
Published: (2023)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
by: Salamatian, Ali, et al.
Published: (2026)
by: Salamatian, Ali, et al.
Published: (2026)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
by: Rahman, Tanzila, et al.
Published: (2024)
by: Rahman, Tanzila, et al.
Published: (2024)
Framework-agnostic Semantically-aware Global Reasoning for Segmentation
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2022)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2022)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video Streams
by: Wu, Zike, et al.
Published: (2025)
by: Wu, Zike, et al.
Published: (2025)
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
by: Salamatian, Ali, et al.
Published: (2025)
by: Salamatian, Ali, et al.
Published: (2025)
One Adapter for All: Towards Unified Representation in Step-Imbalanced Class-Incremental Learning
by: Zhang, Xiaoyan, et al.
Published: (2026)
by: Zhang, Xiaoyan, et al.
Published: (2026)
One Framework to Rule Them All: Unifying RL-Based and RL-Free Methods in RLHF
by: Cai, Xin
Published: (2025)
by: Cai, Xin
Published: (2025)
Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
by: Xu, Bicheng, et al.
Published: (2024)
by: Xu, Bicheng, et al.
Published: (2024)
SIG: A Synthetic Identity Generation Pipeline for Generating Evaluation Datasets for Face Recognition
by: Nzalasse, Kassi, et al.
Published: (2024)
by: Nzalasse, Kassi, et al.
Published: (2024)
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
What Has Been Overlooked in Contrastive Source-Free Domain Adaptation: Leveraging Source-Informed Latent Augmentation within Neighborhood Context
by: Wang, Jing, et al.
Published: (2024)
by: Wang, Jing, et al.
Published: (2024)
MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
by: Awadalla, Anas, et al.
Published: (2024)
by: Awadalla, Anas, et al.
Published: (2024)
MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillation
by: Fu, Yuxiang, et al.
Published: (2025)
by: Fu, Yuxiang, et al.
Published: (2025)
Visual Car Brand Classification by Implementing a Synthetic Image Dataset Creation Pipeline
by: Lippemeier, Jan, et al.
Published: (2024)
by: Lippemeier, Jan, et al.
Published: (2024)
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
by: Luo, Jiayun, et al.
Published: (2024)
by: Luo, Jiayun, et al.
Published: (2024)
Unified Multimodal Uncertain Inference
by: Zhang, Dengjia, et al.
Published: (2026)
by: Zhang, Dengjia, et al.
Published: (2026)
One Noise to Rule Them All: Learning a Unified Model of Spatially-Varying Noise Patterns
by: Maesumi, Arman, et al.
Published: (2024)
by: Maesumi, Arman, et al.
Published: (2024)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
by: Zhou, Zhongyi, et al.
Published: (2025)
by: Zhou, Zhongyi, et al.
Published: (2025)
Harmonizing the Deep: A Unified Information Pipeline for Robust Marine Biodiversity Assessment Across Heterogeneous Domains
by: Piccolo, Marco, et al.
Published: (2026)
by: Piccolo, Marco, et al.
Published: (2026)
UL-DD: A Multimodal Drowsiness Dataset Using Video, Biometric Signals, and Behavioral Data
by: Bodaghi, Morteza, et al.
Published: (2025)
by: Bodaghi, Morteza, et al.
Published: (2025)
SymmetricDiffusers: Learning Discrete Diffusion on Finite Symmetric Groups
by: Zhang, Yongxing, et al.
Published: (2024)
by: Zhang, Yongxing, et al.
Published: (2024)
A Modular Zero-Shot Pipeline for Accident Detection, Localization, and Classification in Traffic Surveillance Video
by: Thakur, Amey, et al.
Published: (2026)
by: Thakur, Amey, et al.
Published: (2026)
One-shot Federated Learning via Synthetic Distiller-Distillate Communication
by: Zhang, Junyuan, et al.
Published: (2024)
by: Zhang, Junyuan, et al.
Published: (2024)
Semantic Residual for Multimodal Unified Discrete Representation
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
Synthetically Enhanced: Unveiling Synthetic Data's Potential in Medical Imaging Research
by: Khosravi, Bardia, et al.
Published: (2023)
by: Khosravi, Bardia, et al.
Published: (2023)
Do Understanding and Generation Fight? A Diagnostic Study of DPO for Unified Multimodal Models
by: Rao, Abinav, et al.
Published: (2026)
by: Rao, Abinav, et al.
Published: (2026)
Online Video Understanding: OVBench and VideoChat-Online
by: Huang, Zhenpeng, et al.
Published: (2024)
by: Huang, Zhenpeng, et al.
Published: (2024)
Understanding Imbalanced Forgetting in Rehearsal-Based Class-Incremental Learning
by: Tamajo, Alberto, et al.
Published: (2026)
by: Tamajo, Alberto, et al.
Published: (2026)
Synthetic Melanoma Image Generation and Evaluation Using Generative Adversarial Networks
by: Lin, Pei-Yu, et al.
Published: (2026)
by: Lin, Pei-Yu, et al.
Published: (2026)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
by: Luo, Jiayun, et al.
Published: (2025)
by: Luo, Jiayun, et al.
Published: (2025)
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
by: Mahdizadeh, Ailar, et al.
Published: (2026)
by: Mahdizadeh, Ailar, et al.
Published: (2026)
Supervised Contrastive Frame Aggregation for Video Representation Learning
by: Chowdhury, Shaif, et al.
Published: (2025)
by: Chowdhury, Shaif, et al.
Published: (2025)
Exploring the Potential of Synthetic Data to Replace Real Data
by: Lee, Hyungtae, et al.
Published: (2024)
by: Lee, Hyungtae, et al.
Published: (2024)
ReXInTheWild: A Unified Benchmark for Medical Photograph Understanding
by: Banerjee, Oishi, et al.
Published: (2026)
by: Banerjee, Oishi, et al.
Published: (2026)
Similar Items
-
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
by: Fan, Wan-Cyuan, et al.
Published: (2024) -
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
by: Bhatt, Gaurav, et al.
Published: (2024) -
Unified Classification and Rejection: A One-versus-All Framework
by: Cheng, Zhen, et al.
Published: (2023) -
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
by: Salamatian, Ali, et al.
Published: (2026) -
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
by: Rahman, Tanzila, et al.
Published: (2024)