Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Haoxin, Li, Boyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
by: Du, Zilin, et al.
Published: (2024)
by: Du, Zilin, et al.
Published: (2024)
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
by: Mishra, Samarth, et al.
Published: (2025)
by: Mishra, Samarth, et al.
Published: (2025)
Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
by: Li, Chenjun, et al.
Published: (2025)
by: Li, Chenjun, et al.
Published: (2025)
Concept-skill Transferability-based Data Selection for Large Vision-Language Models
by: Lee, Jaewoo, et al.
Published: (2024)
by: Lee, Jaewoo, et al.
Published: (2024)
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
Learning to Animate Images from A Few Videos to Portray Delicate Human Actions
by: Li, Haoxin, et al.
Published: (2025)
by: Li, Haoxin, et al.
Published: (2025)
AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks
by: Li, You, et al.
Published: (2024)
by: Li, You, et al.
Published: (2024)
LM4LV: A Frozen Large Language Model for Low-level Vision Tasks
by: Zheng, Boyang, et al.
Published: (2024)
by: Zheng, Boyang, et al.
Published: (2024)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
by: Vani, Sameep, et al.
Published: (2025)
by: Vani, Sameep, et al.
Published: (2025)
Vision-Language Synthetic Data Enhances Echocardiography Downstream Tasks
by: Ashrafian, Pooria, et al.
Published: (2024)
by: Ashrafian, Pooria, et al.
Published: (2024)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
by: Truong, Thanh-Dat, et al.
Published: (2025)
by: Truong, Thanh-Dat, et al.
Published: (2025)
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
by: Ogezi, Michael, et al.
Published: (2025)
by: Ogezi, Michael, et al.
Published: (2025)
Understanding Long Videos with Multimodal Language Models
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
3D Vision and Language Pretraining with Large-Scale Synthetic Data
by: Yang, Dejie, et al.
Published: (2024)
by: Yang, Dejie, et al.
Published: (2024)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
by: Ma, Jingtian, et al.
Published: (2025)
by: Ma, Jingtian, et al.
Published: (2025)
Vision-Enhanced Large Language Models for High-Resolution Image Synthesis and Multimodal Data Interpretation
by: KV, Karthikeya
Published: (2025)
by: KV, Karthikeya
Published: (2025)
Inference-Time Structural Reasoning for Compositional Vision-Language Understanding
by: Bhattacharya, Amartya
Published: (2026)
by: Bhattacharya, Amartya
Published: (2026)
Train a Unified Multimodal Data Quality Classifier with Synthetic Data
by: Wang, Weizhi, et al.
Published: (2025)
by: Wang, Weizhi, et al.
Published: (2025)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
Understanding Degradation with Vision Language Model
by: Lan, Guanzhou, et al.
Published: (2026)
by: Lan, Guanzhou, et al.
Published: (2026)
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models
by: Zhang, Zhifang, et al.
Published: (2025)
by: Zhang, Zhifang, et al.
Published: (2025)
Mema: Memory-Augmented Adapter for Enhanced Vision-Language Understanding
by: Liu, Ying, et al.
Published: (2026)
by: Liu, Ying, et al.
Published: (2026)
Dynamic Multimodal Prototype Learning in Vision-Language Models
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction
by: Chen, Jun, et al.
Published: (2022)
by: Chen, Jun, et al.
Published: (2022)
Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models
by: Guo, Boyang, et al.
Published: (2026)
by: Guo, Boyang, et al.
Published: (2026)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
by: Zhang, Da, et al.
Published: (2025)
by: Zhang, Da, et al.
Published: (2025)
Stable Diffusion for Data Augmentation in COCO and Weed Datasets
by: Deng, Boyang
Published: (2023)
by: Deng, Boyang
Published: (2023)
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
by: Hannan, Tanveer, et al.
Published: (2025)
by: Hannan, Tanveer, et al.
Published: (2025)
MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
Scaling Up Forest Vision with Synthetic Data
by: She, Yihang, et al.
Published: (2025)
by: She, Yihang, et al.
Published: (2025)
Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models
by: Luo, Jiayun, et al.
Published: (2023)
by: Luo, Jiayun, et al.
Published: (2023)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
Rethinking the Mixture of Vision Encoders Paradigm for Enhanced Visual Understanding in Multimodal LLMs
by: Azadani, Mozhgan Nasr, et al.
Published: (2025)
by: Azadani, Mozhgan Nasr, et al.
Published: (2025)
Compositional Kronecker Context Optimization for Vision-Language Models
by: Ding, Kun, et al.
Published: (2024)
by: Ding, Kun, et al.
Published: (2024)
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024)
by: Wijaya, Robert, et al.
Published: (2024)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
by: Rahman, Tanzila, et al.
Published: (2026)
by: Rahman, Tanzila, et al.
Published: (2026)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
Similar Items
-
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
by: Du, Zilin, et al.
Published: (2024) -
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
by: Mishra, Samarth, et al.
Published: (2025) -
Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
by: Li, Chenjun, et al.
Published: (2025) -
Concept-skill Transferability-based Data Selection for Large Vision-Language Models
by: Lee, Jaewoo, et al.
Published: (2024) -
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
by: Kelly, Chris, et al.
Published: (2024)