ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Panwang, Zhao, Jingjing, Lin, Yuchen, Lin, Chenguo, Li, Chenxin, Liu, Hengyu, Shen, Tingting, MU, Yadong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
by: Lin, Yuchen, et al.
Published: (2025)
by: Lin, Yuchen, et al.
Published: (2025)
HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
by: Pan, Panwang, et al.
Published: (2025)
by: Pan, Panwang, et al.
Published: (2025)
InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior
by: Lin, Chenguo, et al.
Published: (2024)
by: Lin, Chenguo, et al.
Published: (2024)
Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
by: Pan, Panwang, et al.
Published: (2025)
by: Pan, Panwang, et al.
Published: (2025)
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
by: Lin, Chenguo, et al.
Published: (2025)
by: Lin, Chenguo, et al.
Published: (2025)
OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation
by: Lin, Yuchen, et al.
Published: (2025)
by: Lin, Yuchen, et al.
Published: (2025)
MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
by: Lin, Chenguo, et al.
Published: (2025)
by: Lin, Chenguo, et al.
Published: (2025)
InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior
by: Lin, Chenguo, et al.
Published: (2024)
by: Lin, Chenguo, et al.
Published: (2024)
HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors
by: Pan, Panwang, et al.
Published: (2024)
by: Pan, Panwang, et al.
Published: (2024)
GaussianStego: A Generalizable Stenography Pipeline for Generative 3D Gaussians Splatting
by: Li, Chenxin, et al.
Published: (2024)
by: Li, Chenxin, et al.
Published: (2024)
NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
by: Xu, Peiran, et al.
Published: (2025)
by: Xu, Peiran, et al.
Published: (2025)
CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration
by: Lin, Yunlong, et al.
Published: (2025)
by: Lin, Yunlong, et al.
Published: (2025)
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
by: Liu, Yaofang, et al.
Published: (2023)
by: Liu, Yaofang, et al.
Published: (2023)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
by: Yan, Yunzhi, et al.
Published: (2024)
by: Yan, Yunzhi, et al.
Published: (2024)
AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation
by: Xu, Ziyi, et al.
Published: (2024)
by: Xu, Ziyi, et al.
Published: (2024)
JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
by: Lin, Yunlong, et al.
Published: (2025)
by: Lin, Yunlong, et al.
Published: (2025)
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
by: Hu, Wenbo, et al.
Published: (2024)
by: Hu, Wenbo, et al.
Published: (2024)
StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter
by: Liu, Gongye, et al.
Published: (2023)
by: Liu, Gongye, et al.
Published: (2023)
ToonCrafter: Generative Cartoon Interpolation
by: Xing, Jinbo, et al.
Published: (2024)
by: Xing, Jinbo, et al.
Published: (2024)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
Endora: Video Generation Models as Endoscopy Simulators
by: Li, Chenxin, et al.
Published: (2024)
by: Li, Chenxin, et al.
Published: (2024)
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
by: Lin, Jiawen, et al.
Published: (2025)
by: Lin, Jiawen, et al.
Published: (2025)
PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
by: Mao, Qing, et al.
Published: (2025)
by: Mao, Qing, et al.
Published: (2025)
DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
by: Wen, Kairun, et al.
Published: (2025)
by: Wen, Kairun, et al.
Published: (2025)
MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
by: Zheng, Lihao, et al.
Published: (2025)
by: Zheng, Lihao, et al.
Published: (2025)
MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation
by: Zhao, Yuan, et al.
Published: (2025)
by: Zhao, Yuan, et al.
Published: (2025)
MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
by: Ling, Run, et al.
Published: (2025)
by: Ling, Run, et al.
Published: (2025)
Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
by: Huang, Yuzhi, et al.
Published: (2025)
by: Huang, Yuzhi, et al.
Published: (2025)
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
by: Chen, Haoxin, et al.
Published: (2024)
by: Chen, Haoxin, et al.
Published: (2024)
PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control
by: Zhong, Yong, et al.
Published: (2024)
by: Zhong, Yong, et al.
Published: (2024)
Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
by: Zhang, Guiyu, et al.
Published: (2025)
by: Zhang, Guiyu, et al.
Published: (2025)
EmotiCrafter: Text-to-Emotional-Image Generation based on Valence-Arousal Model
by: Dang, Shengqi, et al.
Published: (2025)
by: Dang, Shengqi, et al.
Published: (2025)
CityCraft: A Real Crafter for 3D City Generation
by: Deng, Jie, et al.
Published: (2024)
by: Deng, Jie, et al.
Published: (2024)
SceneCrafter: Controllable Multi-View Driving Scene Editing
by: Zhu, Zehao, et al.
Published: (2025)
by: Zhu, Zehao, et al.
Published: (2025)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
by: Huang, Kaiyi, et al.
Published: (2024)
by: Huang, Kaiyi, et al.
Published: (2024)
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
by: Zhao, Haozhe, et al.
Published: (2026)
by: Zhao, Haozhe, et al.
Published: (2026)
NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors
by: Bin, Yanrui, et al.
Published: (2025)
by: Bin, Yanrui, et al.
Published: (2025)
Similar Items
-
PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
by: Lin, Yuchen, et al.
Published: (2025) -
HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
by: Pan, Panwang, et al.
Published: (2025) -
InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior
by: Lin, Chenguo, et al.
Published: (2024) -
Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
by: Pan, Panwang, et al.
Published: (2025) -
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
by: Lin, Chenguo, et al.
Published: (2025)