MV-S2V: Multi-View Subject-Consistent Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Ziyang, Gong, Xinyu, Liu, Bangya, Zhao, Zelin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoMV: Consistent Multi-View Generation Based on Large Video Generative Model
by: Zuo, Qi, et al.
Published: (2024)
by: Zuo, Qi, et al.
Published: (2024)
ByteLoom: Weaving Geometry-Consistent Human-Object Interactions through Progressive Curriculum Learning
by: Liu, Bangya, et al.
Published: (2025)
by: Liu, Bangya, et al.
Published: (2025)
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
by: Hu, Wenbo, et al.
Published: (2024)
by: Hu, Wenbo, et al.
Published: (2024)
V-LASIK: Consistent Glasses-Removal from Videos Using Synthetic Data
by: Shalev-Arkushin, Rotem, et al.
Published: (2024)
by: Shalev-Arkushin, Rotem, et al.
Published: (2024)
GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
by: Xu, Tian-Xing, et al.
Published: (2025)
by: Xu, Tian-Xing, et al.
Published: (2025)
MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars
by: Taubner, Felix, et al.
Published: (2025)
by: Taubner, Felix, et al.
Published: (2025)
3DPhysVideo: Consistency-Guided Flow SDE for Video Generation via 3D Scene Reconstruction and Physical Simulation
by: Kim, Hwidong, et al.
Published: (2026)
by: Kim, Hwidong, et al.
Published: (2026)
Generative Object Insertion in Gaussian Splatting with a Multi-View Diffusion Model
by: Zhong, Hongliang, et al.
Published: (2024)
by: Zhong, Hongliang, et al.
Published: (2024)
Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
by: Song, Chenxi, et al.
Published: (2025)
by: Song, Chenxi, et al.
Published: (2025)
ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
by: Liu, Fangfu, et al.
Published: (2024)
by: Liu, Fangfu, et al.
Published: (2024)
Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
by: Gu, Zekai, et al.
Published: (2025)
by: Gu, Zekai, et al.
Published: (2025)
MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation
by: Wang, Kuan-Chieh, et al.
Published: (2024)
by: Wang, Kuan-Chieh, et al.
Published: (2024)
MVTN: Learning Multi-View Transformations for 3D Understanding
by: Hamdi, Abdullah, et al.
Published: (2022)
by: Hamdi, Abdullah, et al.
Published: (2022)
GraphicsDreamer: Image to 3D Generation with Physical Consistency
by: Chen, Pei, et al.
Published: (2024)
by: Chen, Pei, et al.
Published: (2024)
Be Decisive: Noise-Induced Layouts for Multi-Subject Generation
by: Dahary, Omer, et al.
Published: (2025)
by: Dahary, Omer, et al.
Published: (2025)
VideoPanda: Video Panoramic Diffusion with Multi-view Attention
by: Xie, Kevin, et al.
Published: (2025)
by: Xie, Kevin, et al.
Published: (2025)
DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware Diffusion
by: Ye, Weicai, et al.
Published: (2024)
by: Ye, Weicai, et al.
Published: (2024)
ObjectMover: Generative Object Movement with Video Prior
by: Yu, Xin, et al.
Published: (2025)
by: Yu, Xin, et al.
Published: (2025)
Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation
by: Dahary, Omer, et al.
Published: (2024)
by: Dahary, Omer, et al.
Published: (2024)
Physical Simulator In-the-Loop Video Generation
by: Foo, Lin Geng, et al.
Published: (2026)
by: Foo, Lin Geng, et al.
Published: (2026)
Mixture of Contexts for Long Video Generation
by: Cai, Shengqu, et al.
Published: (2025)
by: Cai, Shengqu, et al.
Published: (2025)
Improving Video Generation with Human Feedback
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
RealWonder: Real-Time Physical Action-Conditioned Video Generation
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
by: Jiang, Liyao, et al.
Published: (2024)
by: Jiang, Liyao, et al.
Published: (2024)
LatentEdit: Adaptive Latent Control for Consistent Semantic Editing
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
DreamDrive: Generative 4D Scene Modeling from Street View Images
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
Bridging Text and Video Generation: A Survey
by: Kumar, Nilay, et al.
Published: (2025)
by: Kumar, Nilay, et al.
Published: (2025)
AMG: Avatar Motion Guided Video Generation
by: Yang, Zhangsihao, et al.
Published: (2024)
by: Yang, Zhangsihao, et al.
Published: (2024)
MeshAnything V2: Artist-Created Mesh Generation With Adjacent Mesh Tokenization
by: Chen, Yiwen, et al.
Published: (2024)
by: Chen, Yiwen, et al.
Published: (2024)
PhysGaia: A Physics-Aware Benchmark with Multi-Body Interactions for Dynamic Novel View Synthesis
by: Kim, Mijeong, et al.
Published: (2025)
by: Kim, Mijeong, et al.
Published: (2025)
Evaluating Design Video Generation: Metrics for Compositional Fidelity
by: Deganutti, Adrienne, et al.
Published: (2026)
by: Deganutti, Adrienne, et al.
Published: (2026)
SeqTex: Generate Mesh Textures in Video Sequence
by: Yuan, Ze, et al.
Published: (2025)
by: Yuan, Ze, et al.
Published: (2025)
GarmentCrafter: Progressive Novel View Synthesis for Single-View 3D Garment Reconstruction and Editing
by: Wang, Yuanhao, et al.
Published: (2025)
by: Wang, Yuanhao, et al.
Published: (2025)
MotionV2V: Editing Motion in a Video
by: Burgert, Ryan, et al.
Published: (2025)
by: Burgert, Ryan, et al.
Published: (2025)
SpaRP: Fast 3D Object Reconstruction and Pose Estimation from Sparse Views
by: Xu, Chao, et al.
Published: (2024)
by: Xu, Chao, et al.
Published: (2024)
City-Mesh3R: Simulation-Ready City-Scale 3D Mesh Reconstruction from Multi-View Images
by: Paul, Sayan, et al.
Published: (2026)
by: Paul, Sayan, et al.
Published: (2026)
Oitijjo-3D: Generative AI Framework for Rapid 3D Heritage Reconstruction from Street View Imagery
by: Ope, Momen Khandoker, et al.
Published: (2025)
by: Ope, Momen Khandoker, et al.
Published: (2025)
GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion
by: Tang, Jiapeng, et al.
Published: (2024)
by: Tang, Jiapeng, et al.
Published: (2024)
Spatiotemporally Consistent Indoor Lighting Estimation with Diffusion Priors
by: Tong, Mutian, et al.
Published: (2025)
by: Tong, Mutian, et al.
Published: (2025)
VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping
by: Shao, Hao, et al.
Published: (2024)
by: Shao, Hao, et al.
Published: (2024)
Similar Items
-
VideoMV: Consistent Multi-View Generation Based on Large Video Generative Model
by: Zuo, Qi, et al.
Published: (2024) -
ByteLoom: Weaving Geometry-Consistent Human-Object Interactions through Progressive Curriculum Learning
by: Liu, Bangya, et al.
Published: (2025) -
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
by: Hu, Wenbo, et al.
Published: (2024) -
V-LASIK: Consistent Glasses-Removal from Videos Using Synthetic Data
by: Shalev-Arkushin, Rotem, et al.
Published: (2024) -
GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
by: Xu, Tian-Xing, et al.
Published: (2025)