ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Pan, Panwang, Zhao, Jingjing, Lin, Yuchen, Lin, Chenguo, Li, Chenxin, Liu, Hengyu, Shen, Tingting, MU, Yadong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
di: Lin, Yuchen, et al.
Pubblicazione: (2025)
di: Lin, Yuchen, et al.
Pubblicazione: (2025)
HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
di: Pan, Panwang, et al.
Pubblicazione: (2025)
di: Pan, Panwang, et al.
Pubblicazione: (2025)
InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior
di: Lin, Chenguo, et al.
Pubblicazione: (2024)
di: Lin, Chenguo, et al.
Pubblicazione: (2024)
Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
di: Pan, Panwang, et al.
Pubblicazione: (2025)
di: Pan, Panwang, et al.
Pubblicazione: (2025)
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
di: Lin, Chenguo, et al.
Pubblicazione: (2025)
di: Lin, Chenguo, et al.
Pubblicazione: (2025)
OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation
di: Lin, Yuchen, et al.
Pubblicazione: (2025)
di: Lin, Yuchen, et al.
Pubblicazione: (2025)
MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
di: Lin, Chenguo, et al.
Pubblicazione: (2025)
di: Lin, Chenguo, et al.
Pubblicazione: (2025)
InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior
di: Lin, Chenguo, et al.
Pubblicazione: (2024)
di: Lin, Chenguo, et al.
Pubblicazione: (2024)
HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors
di: Pan, Panwang, et al.
Pubblicazione: (2024)
di: Pan, Panwang, et al.
Pubblicazione: (2024)
GaussianStego: A Generalizable Stenography Pipeline for Generative 3D Gaussians Splatting
di: Li, Chenxin, et al.
Pubblicazione: (2024)
di: Li, Chenxin, et al.
Pubblicazione: (2024)
NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
di: Xu, Peiran, et al.
Pubblicazione: (2025)
di: Xu, Peiran, et al.
Pubblicazione: (2025)
CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
di: Wu, Tao, et al.
Pubblicazione: (2024)
di: Wu, Tao, et al.
Pubblicazione: (2024)
MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
di: Wu, Tao, et al.
Pubblicazione: (2025)
di: Wu, Tao, et al.
Pubblicazione: (2025)
JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration
di: Lin, Yunlong, et al.
Pubblicazione: (2025)
di: Lin, Yunlong, et al.
Pubblicazione: (2025)
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
di: Liu, Yaofang, et al.
Pubblicazione: (2023)
di: Liu, Yaofang, et al.
Pubblicazione: (2023)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
di: Yan, Yunzhi, et al.
Pubblicazione: (2024)
di: Yan, Yunzhi, et al.
Pubblicazione: (2024)
AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation
di: Xu, Ziyi, et al.
Pubblicazione: (2024)
di: Xu, Ziyi, et al.
Pubblicazione: (2024)
JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
di: Lin, Yunlong, et al.
Pubblicazione: (2025)
di: Lin, Yunlong, et al.
Pubblicazione: (2025)
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter
di: Liu, Gongye, et al.
Pubblicazione: (2023)
di: Liu, Gongye, et al.
Pubblicazione: (2023)
ToonCrafter: Generative Cartoon Interpolation
di: Xing, Jinbo, et al.
Pubblicazione: (2024)
di: Xing, Jinbo, et al.
Pubblicazione: (2024)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
di: Xu, Runsen, et al.
Pubblicazione: (2024)
di: Xu, Runsen, et al.
Pubblicazione: (2024)
Endora: Video Generation Models as Endoscopy Simulators
di: Li, Chenxin, et al.
Pubblicazione: (2024)
di: Li, Chenxin, et al.
Pubblicazione: (2024)
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
di: Lin, Jiawen, et al.
Pubblicazione: (2025)
di: Lin, Jiawen, et al.
Pubblicazione: (2025)
PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
di: Mao, Qing, et al.
Pubblicazione: (2025)
di: Mao, Qing, et al.
Pubblicazione: (2025)
DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
di: Wen, Kairun, et al.
Pubblicazione: (2025)
di: Wen, Kairun, et al.
Pubblicazione: (2025)
MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
di: Zheng, Lihao, et al.
Pubblicazione: (2025)
di: Zheng, Lihao, et al.
Pubblicazione: (2025)
MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation
di: Zhao, Yuan, et al.
Pubblicazione: (2025)
di: Zhao, Yuan, et al.
Pubblicazione: (2025)
MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
di: Ling, Run, et al.
Pubblicazione: (2025)
di: Ling, Run, et al.
Pubblicazione: (2025)
Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
di: Huang, Yuzhi, et al.
Pubblicazione: (2025)
di: Huang, Yuzhi, et al.
Pubblicazione: (2025)
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
di: Chen, Haoxin, et al.
Pubblicazione: (2024)
di: Chen, Haoxin, et al.
Pubblicazione: (2024)
PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control
di: Zhong, Yong, et al.
Pubblicazione: (2024)
di: Zhong, Yong, et al.
Pubblicazione: (2024)
Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
di: Zhang, Guiyu, et al.
Pubblicazione: (2025)
di: Zhang, Guiyu, et al.
Pubblicazione: (2025)
EmotiCrafter: Text-to-Emotional-Image Generation based on Valence-Arousal Model
di: Dang, Shengqi, et al.
Pubblicazione: (2025)
di: Dang, Shengqi, et al.
Pubblicazione: (2025)
CityCraft: A Real Crafter for 3D City Generation
di: Deng, Jie, et al.
Pubblicazione: (2024)
di: Deng, Jie, et al.
Pubblicazione: (2024)
SceneCrafter: Controllable Multi-View Driving Scene Editing
di: Zhu, Zehao, et al.
Pubblicazione: (2025)
di: Zhu, Zehao, et al.
Pubblicazione: (2025)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
di: Huang, Kaiyi, et al.
Pubblicazione: (2024)
di: Huang, Kaiyi, et al.
Pubblicazione: (2024)
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
di: Zhao, Haozhe, et al.
Pubblicazione: (2026)
di: Zhao, Haozhe, et al.
Pubblicazione: (2026)
NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors
di: Bin, Yanrui, et al.
Pubblicazione: (2025)
di: Bin, Yanrui, et al.
Pubblicazione: (2025)
Documenti analoghi
-
PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
di: Lin, Yuchen, et al.
Pubblicazione: (2025) -
HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
di: Pan, Panwang, et al.
Pubblicazione: (2025) -
InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior
di: Lin, Chenguo, et al.
Pubblicazione: (2024) -
Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
di: Pan, Panwang, et al.
Pubblicazione: (2025) -
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
di: Lin, Chenguo, et al.
Pubblicazione: (2025)