World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jiacong, Wu, Bohong, Jiang, Haiyong, Zhou, Xun, Xiao, Xin, Guo, Haoyuan, Xiao, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
by: Xiao, Xin, et al.
Published: (2024)
by: Xiao, Xin, et al.
Published: (2024)
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024)
by: Dong, Hongyuan, et al.
Published: (2024)
VGR: Visual Grounded Reasoning
by: Wang, Jiacong, et al.
Published: (2025)
by: Wang, Jiacong, et al.
Published: (2025)
DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism
by: Wang, Zhen, et al.
Published: (2023)
by: Wang, Zhen, et al.
Published: (2023)
Instruct-Imagen: Image Generation with Multi-modal Instruction
by: Hu, Hexiang, et al.
Published: (2024)
by: Hu, Hexiang, et al.
Published: (2024)
HierRelTriple: Guiding Indoor Layout Generation with Hierarchical Relationship Triplet Losses
by: Sun, Kaifan, et al.
Published: (2025)
by: Sun, Kaifan, et al.
Published: (2025)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
by: Xie, Senwei, et al.
Published: (2025)
by: Xie, Senwei, et al.
Published: (2025)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
PS-CAD: Local Geometry Guidance via Prompting and Selection for CAD Reconstruction
by: Yang, Bingchen, et al.
Published: (2024)
by: Yang, Bingchen, et al.
Published: (2024)
SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part Segmentation
by: Hu, Yueyang, et al.
Published: (2025)
by: Hu, Yueyang, et al.
Published: (2025)
Empowering Vector Graphics with Consistently Arbitrary Viewing and View-dependent Visibility
by: Li, Yidi, et al.
Published: (2025)
by: Li, Yidi, et al.
Published: (2025)
DyStream: Streaming Dyadic Talking Heads Generation via Flow Matching-based Autoregressive Model
by: Chen, Bohong, et al.
Published: (2025)
by: Chen, Bohong, et al.
Published: (2025)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
by: Huang, Junming, et al.
Published: (2026)
by: Huang, Junming, et al.
Published: (2026)
Visual Object Tracking on Multi-modal RGB-D Videos: A Review
by: Zhu, Xue-Feng, et al.
Published: (2022)
by: Zhu, Xue-Feng, et al.
Published: (2022)
InstructMoLE: Instruction-Guided Mixture of Low-rank Experts for Multi-Conditional Image Generation
by: Xiao, Jinqi, et al.
Published: (2025)
by: Xiao, Jinqi, et al.
Published: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
Explainable, Multi-modal Wound Infection Classification from Images Augmented with Generated Captions
by: Busaranuvong, Palawat, et al.
Published: (2025)
by: Busaranuvong, Palawat, et al.
Published: (2025)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
by: Safaei, Bardia, et al.
Published: (2025)
by: Safaei, Bardia, et al.
Published: (2025)
CoMA: Compositional Human Motion Generation with Multi-modal Agents
by: Sun, Shanlin, et al.
Published: (2024)
by: Sun, Shanlin, et al.
Published: (2024)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
by: Jiang, Lingjie, et al.
Published: (2025)
by: Jiang, Lingjie, et al.
Published: (2025)
Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training
by: Xing, Jinbo, et al.
Published: (2026)
by: Xing, Jinbo, et al.
Published: (2026)
SCA3D: Enhancing Cross-modal 3D Retrieval via 3D Shape and Caption Paired Data Augmentation
by: Ren, Junlong, et al.
Published: (2025)
by: Ren, Junlong, et al.
Published: (2025)
Endless World: Real-Time 3D-Aware Long Video Generation
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
by: Sakai, Yuki, et al.
Published: (2025)
by: Sakai, Yuki, et al.
Published: (2025)
PhotoFramer: Multi-modal Image Composition Instruction
by: You, Zhiyuan, et al.
Published: (2025)
by: You, Zhiyuan, et al.
Published: (2025)
Ev-Layout: A Large-scale Event-based Multi-modal Dataset for Indoor Layout Estimation and Tracking
by: Guo, Xucheng, et al.
Published: (2025)
by: Guo, Xucheng, et al.
Published: (2025)
InstructSAM: Segment Any Instance with Any Instructions
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation
by: Qin, Yiran, et al.
Published: (2026)
by: Qin, Yiran, et al.
Published: (2026)
GDDS: A Single Domain Generalized Defect Detection Frame of Open World Scenario using Gather and Distribute Domain-shift Suppression Network
by: Chen, Haiyong, et al.
Published: (2024)
by: Chen, Haiyong, et al.
Published: (2024)
Multi-modal Attribute Prompting for Vision-Language Models
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning
by: Zhou, Qifeng, et al.
Published: (2024)
by: Zhou, Qifeng, et al.
Published: (2024)
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
by: Qiu, Xuerui, et al.
Published: (2026)
by: Qiu, Xuerui, et al.
Published: (2026)
Enhanced Masked Image Modeling to Avoid Model Collapse on Multi-modal MRI Datasets
by: Han, Linxuan, et al.
Published: (2024)
by: Han, Linxuan, et al.
Published: (2024)
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
by: Wu, Chengyue, et al.
Published: (2024)
by: Wu, Chengyue, et al.
Published: (2024)
Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-learning
by: Xie, Zhuyang, et al.
Published: (2024)
by: Xie, Zhuyang, et al.
Published: (2024)
Unveiling the Tapestry of Consistency in Large Vision-Language Models
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
Similar Items
-
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
by: Xiao, Xin, et al.
Published: (2024) -
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024) -
VGR: Visual Grounded Reasoning
by: Wang, Jiacong, et al.
Published: (2025) -
DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism
by: Wang, Zhen, et al.
Published: (2023) -
Instruct-Imagen: Image Generation with Multi-modal Instruction
by: Hu, Hexiang, et al.
Published: (2024)