Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition
Fuente:
arXiv
Saved in:
| Main Authors: | Yeh, Chun-Hsiao, Cheng, Ta-Ying, Hsieh, He-Yen, Lin, Chuan-En, Ma, Yi, Markham, Andrew, Trigoni, Niki, Kung, H. T., Chen, Yubei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
by: Ma, Chenyang, et al.
Published: (2024)
by: Ma, Chenyang, et al.
Published: (2024)
ZeST: Zero-Shot Material Transfer from a Single Image
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
Learning Continuous 3D Words for Text-to-Image Generation
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
Pre-training Feature Guided Diffusion Model for Speech Enhancement
by: Yang, Yiyuan, et al.
Published: (2024)
by: Yang, Yiyuan, et al.
Published: (2024)
WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization
by: Wang, Jialu, et al.
Published: (2024)
by: Wang, Jialu, et al.
Published: (2024)
Mitigating Cognitive Bias in RLHF by Altering Rationality
by: Horter, Tiffany, et al.
Published: (2026)
by: Horter, Tiffany, et al.
Published: (2026)
MambaLoc: Efficient Camera Localisation via State Space Model
by: Wang, Jialu, et al.
Published: (2024)
by: Wang, Jialu, et al.
Published: (2024)
Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
by: Xu, Shitong, et al.
Published: (2025)
by: Xu, Shitong, et al.
Published: (2025)
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
by: Yang, Yiyuan, et al.
Published: (2025)
by: Yang, Yiyuan, et al.
Published: (2025)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
Data Factory with Minimal Human Effort Using VLMs
by: Ye, Jiaojiao, et al.
Published: (2025)
by: Ye, Jiaojiao, et al.
Published: (2025)
Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical Representation
by: Shin, Sangyun, et al.
Published: (2023)
by: Shin, Sangyun, et al.
Published: (2023)
GazeGen: Gaze-Driven User Interaction for Visual Content Generation
by: Hsieh, He-Yen, et al.
Published: (2024)
by: Hsieh, He-Yen, et al.
Published: (2024)
COOPERA: Continual Open-Ended Human-Robot Assistance
by: Ma, Chenyang, et al.
Published: (2025)
by: Ma, Chenyang, et al.
Published: (2025)
SPEAR: Receiver-to-Receiver Acoustic Neural Warping Field
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
Insight: A Multi-Modal Diagnostic Pipeline using LLMs for Ocular Surface Disease Diagnosis
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
Dusk Till Dawn: Self-supervised Nighttime Stereo Depth Estimation using Visual Foundation Models
by: Vankadari, Madhu, et al.
Published: (2024)
by: Vankadari, Madhu, et al.
Published: (2024)
CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
by: Ma, Chenyang, et al.
Published: (2026)
by: Ma, Chenyang, et al.
Published: (2026)
MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation
by: Lan, Yun-Han, et al.
Published: (2024)
by: Lan, Yun-Han, et al.
Published: (2024)
Gen-n-Val: Agentic Image Data Generation and Validation
by: Huang, Jing-En, et al.
Published: (2025)
by: Huang, Jing-En, et al.
Published: (2025)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
GenCompositor: Generative Video Compositing with Diffusion Transformer
by: Yang, Shuzhou, et al.
Published: (2025)
by: Yang, Shuzhou, et al.
Published: (2025)
VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization
by: Zhou, Kaichen, et al.
Published: (2020)
by: Zhou, Kaichen, et al.
Published: (2020)
Towards Multi-Modal Animal Pose Estimation: A Survey and In-Depth Analysis
by: Deng, Qianyi, et al.
Published: (2024)
by: Deng, Qianyi, et al.
Published: (2024)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
by: Zhou, Yifan, et al.
Published: (2026)
by: Zhou, Yifan, et al.
Published: (2026)
DynPoint: Dynamic Neural Point For View Synthesis
by: Zhou, Kaichen, et al.
Published: (2023)
by: Zhou, Kaichen, et al.
Published: (2023)
GenColor: Generative Color-Concept Association in Visual Design
by: Hou, Yihan, et al.
Published: (2025)
by: Hou, Yihan, et al.
Published: (2025)
Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes
by: Zhou, Kaichen, et al.
Published: (2023)
by: Zhou, Kaichen, et al.
Published: (2023)
EDIT: Early Diffusion Inference Termination for dLLMs Based on Dynamics of Training Gradients
by: Hsieh, He-Yen, et al.
Published: (2025)
by: Hsieh, He-Yen, et al.
Published: (2025)
ShapeGen: Robotic Data Generation for Category-Level Manipulation
by: Wang, Yirui, et al.
Published: (2026)
by: Wang, Yirui, et al.
Published: (2026)
EgoGen: An Egocentric Synthetic Data Generator
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
To Search or To Gen? Exploring the Synergy between Generative AI and Web Search in Programming
by: Yen, Ryan, et al.
Published: (2024)
by: Yen, Ryan, et al.
Published: (2024)
SnapMoGen: Human Motion Generation from Expressive Texts
by: Guo, Chuan, et al.
Published: (2025)
by: Guo, Chuan, et al.
Published: (2025)
GenUQ: Predictive Uncertainty Estimates via Generative Hyper-Networks
by: Yen, Tian Yu, et al.
Published: (2025)
by: Yen, Tian Yu, et al.
Published: (2025)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
GenEx: Generating an Explorable World
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
GenDexHand: Generative Simulation for Dexterous Hands
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
GameGen-X: Interactive Open-world Game Video Generation
by: Che, Haoxuan, et al.
Published: (2024)
by: Che, Haoxuan, et al.
Published: (2024)
ArcGen: Generalizing Neural Backdoor Detection Across Diverse Architectures
by: Yang, Zhonghao, et al.
Published: (2025)
by: Yang, Zhonghao, et al.
Published: (2025)
Similar Items
-
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
by: Ma, Chenyang, et al.
Published: (2024) -
ZeST: Zero-Shot Material Transfer from a Single Image
by: Cheng, Ta-Ying, et al.
Published: (2024) -
Learning Continuous 3D Words for Text-to-Image Generation
by: Cheng, Ta-Ying, et al.
Published: (2024) -
Pre-training Feature Guided Diffusion Model for Speech Enhancement
by: Yang, Yiyuan, et al.
Published: (2024) -
WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization
by: Wang, Jialu, et al.
Published: (2024)