MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tong, Jinguang, Wu, Jinbo, Wang, Kaisiyuan, Shen, Zhelun, Huang, Xuan, Xiang, Mochu, Li, Xuesong, Li, Yingying, Feng, Haocheng, Zhao, Chen, Zhou, Hang, He, Wei, Nguyen, Chuong, Wang, Jingdong, Li, Hongdong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer
von: Shen, Zhelun, et al.
Veröffentlicht: (2025)
von: Shen, Zhelun, et al.
Veröffentlicht: (2025)
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
von: Fan, Yingying, et al.
Veröffentlicht: (2025)
von: Fan, Yingying, et al.
Veröffentlicht: (2025)
GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection
von: Huang, Xuan, et al.
Veröffentlicht: (2026)
von: Huang, Xuan, et al.
Veröffentlicht: (2026)
GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction
von: Tong, Jinguang, et al.
Veröffentlicht: (2025)
von: Tong, Jinguang, et al.
Veröffentlicht: (2025)
RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations
von: Xiang, Mochu, et al.
Veröffentlicht: (2026)
von: Xiang, Mochu, et al.
Veröffentlicht: (2026)
TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
von: Guan, Jiazhi, et al.
Veröffentlicht: (2024)
von: Guan, Jiazhi, et al.
Veröffentlicht: (2024)
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
von: Zhang, Chushan, et al.
Veröffentlicht: (2026)
von: Zhang, Chushan, et al.
Veröffentlicht: (2026)
DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary
von: Guan, Jiazhi, et al.
Veröffentlicht: (2026)
von: Guan, Jiazhi, et al.
Veröffentlicht: (2026)
3D-IDE: 3D Implicit Depth Emergent
von: Zhang, Chushan, et al.
Veröffentlicht: (2026)
von: Zhang, Chushan, et al.
Veröffentlicht: (2026)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
von: Sun, Yasheng, et al.
Veröffentlicht: (2025)
von: Sun, Yasheng, et al.
Veröffentlicht: (2025)
Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection
von: Wang, Shan, et al.
Veröffentlicht: (2025)
von: Wang, Shan, et al.
Veröffentlicht: (2025)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
von: Yang, Quanwei, et al.
Veröffentlicht: (2025)
von: Yang, Quanwei, et al.
Veröffentlicht: (2025)
Structural Energy Guidance for View-Consistent Text-to-3D Generation
von: Zhang, Qing, et al.
Veröffentlicht: (2026)
von: Zhang, Qing, et al.
Veröffentlicht: (2026)
Improving Viewpoint Consistency in 3D Generation via Structure Feature and CLIP Guidance
von: Zhang, Qing, et al.
Veröffentlicht: (2024)
von: Zhang, Qing, et al.
Veröffentlicht: (2024)
DGNS: Deformable Gaussian Splatting and Dynamic Neural Surface for Monocular Dynamic 3D Reconstruction
von: Li, Xuesong, et al.
Veröffentlicht: (2024)
von: Li, Xuesong, et al.
Veröffentlicht: (2024)
Structural Energy-Guided Sampling for View-Consistent Text-to-3D
von: Zhang, Qing, et al.
Veröffentlicht: (2025)
von: Zhang, Qing, et al.
Veröffentlicht: (2025)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025)
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025)
Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models
von: Zhang, Qing, et al.
Veröffentlicht: (2026)
von: Zhang, Qing, et al.
Veröffentlicht: (2026)
OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection
von: Hou, Jinghua, et al.
Veröffentlicht: (2024)
von: Hou, Jinghua, et al.
Veröffentlicht: (2024)
DPBridge: Latent Diffusion Bridge for Dense Prediction
von: Ji, Haorui, et al.
Veröffentlicht: (2024)
von: Ji, Haorui, et al.
Veröffentlicht: (2024)
AnyAct: Towards Human Reenactment of Character Motion From Video
von: Chen, Liuhan, et al.
Veröffentlicht: (2026)
von: Chen, Liuhan, et al.
Veröffentlicht: (2026)
Geometry-guided Cross-view Diffusion for One-to-many Cross-view Image Synthesis
von: Lin, Tao Jun, et al.
Veröffentlicht: (2024)
von: Lin, Tao Jun, et al.
Veröffentlicht: (2024)
Enhancing Features in Long-tailed Data Using Large Vision Model
von: Han, Pengxiao, et al.
Veröffentlicht: (2025)
von: Han, Pengxiao, et al.
Veröffentlicht: (2025)
Quantum decision trees with information entropy
von: Li, Zhelun, et al.
Veröffentlicht: (2025)
von: Li, Zhelun, et al.
Veröffentlicht: (2025)
Adaptive and Balanced Re-initialization for Long-timescale Continual Test-time Domain Adaptation
von: Wang, Yanshuo, et al.
Veröffentlicht: (2026)
von: Wang, Yanshuo, et al.
Veröffentlicht: (2026)
Maintain Plasticity in Long-timescale Continual Test-time Adaptation
von: Wang, Yanshuo, et al.
Veröffentlicht: (2024)
von: Wang, Yanshuo, et al.
Veröffentlicht: (2024)
InterDyad: Interactive Dyadic Speech-to-Video Generation by Querying Intermediate Visual Guidance
von: Pan, Dongwei, et al.
Veröffentlicht: (2026)
von: Pan, Dongwei, et al.
Veröffentlicht: (2026)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
von: Guan, Jiazhi, et al.
Veröffentlicht: (2024)
von: Guan, Jiazhi, et al.
Veröffentlicht: (2024)
Monte Carlo Analysis of Boid Simulations with Obstacles: A Physics-Based Perspective
von: Nguyen, Quoc Chuong
Veröffentlicht: (2024)
von: Nguyen, Quoc Chuong
Veröffentlicht: (2024)
Network Sampling: An Overview and Comparative Analysis
von: Nguyen, Quoc Chuong
Veröffentlicht: (2025)
von: Nguyen, Quoc Chuong
Veröffentlicht: (2025)
SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors
von: Li, Aixuan, et al.
Veröffentlicht: (2025)
von: Li, Aixuan, et al.
Veröffentlicht: (2025)
Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection
von: Li, Jiaming, et al.
Veröffentlicht: (2024)
von: Li, Jiaming, et al.
Veröffentlicht: (2024)
Homography Guided Temporal Fusion for Road Line and Marking Segmentation
von: Wang, Shan, et al.
Veröffentlicht: (2024)
von: Wang, Shan, et al.
Veröffentlicht: (2024)
AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation
von: Sun, Yasheng, et al.
Veröffentlicht: (2024)
von: Sun, Yasheng, et al.
Veröffentlicht: (2024)
Contrastive Conditional Latent Diffusion for Audio-visual Segmentation
von: Mao, Yuxin, et al.
Veröffentlicht: (2023)
von: Mao, Yuxin, et al.
Veröffentlicht: (2023)
Bridging Human Interpretation and Machine Representation: A Landscape of Qualitative Data Analysis in the LLM Era
von: Pi, Xinyu, et al.
Veröffentlicht: (2026)
von: Pi, Xinyu, et al.
Veröffentlicht: (2026)
Degeneration of the archimedean height pairing of algebraically trivial cycles
von: Chen, Zhelun
Veröffentlicht: (2025)
von: Chen, Zhelun
Veröffentlicht: (2025)
Anchored Diffusion for Video Face Reenactment
von: Kligvasser, Idan, et al.
Veröffentlicht: (2024)
von: Kligvasser, Idan, et al.
Veröffentlicht: (2024)
WaveletGaussian: Wavelet-domain Diffusion for Sparse-view 3D Gaussian Object Reconstruction
von: Nguyen, Hung, et al.
Veröffentlicht: (2025)
von: Nguyen, Hung, et al.
Veröffentlicht: (2025)
Anchor-free Cross-view Object Geo-localization with Gaussian Position Encoding and Cross-view Association
von: Ling, Xingtao, et al.
Veröffentlicht: (2025)
von: Ling, Xingtao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer
von: Shen, Zhelun, et al.
Veröffentlicht: (2025) -
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
von: Fan, Yingying, et al.
Veröffentlicht: (2025) -
GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection
von: Huang, Xuan, et al.
Veröffentlicht: (2026) -
GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction
von: Tong, Jinguang, et al.
Veröffentlicht: (2025) -
RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations
von: Xiang, Mochu, et al.
Veröffentlicht: (2026)