ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Dong, Mingyu, Xia, Chong, Jia, Mingyuan, Lyu, Weichen, Xu, Long, Zhu, Zheng, Duan, Yueqi |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SimRecon: SimReady Compositional Scene Reconstruction from Real Videos
par: Xia, Chong, et autres
Publié: (2026)
par: Xia, Chong, et autres
Publié: (2026)
ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment
par: Xia, Chong, et autres
Publié: (2025)
par: Xia, Chong, et autres
Publié: (2025)
ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
par: Liu, Fangfu, et autres
Publié: (2024)
par: Liu, Fangfu, et autres
Publié: (2024)
VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
par: Wang, Hanyang, et autres
Publié: (2025)
par: Wang, Hanyang, et autres
Publié: (2025)
GranAlign: Granularity-Aware Alignment Framework for Zero-Shot Video Moment Retrieval
par: Jeon, Mingyu, et autres
Publié: (2026)
par: Jeon, Mingyu, et autres
Publié: (2026)
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
par: Sun, Wenqiang, et autres
Publié: (2024)
par: Sun, Wenqiang, et autres
Publié: (2024)
Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification
par: Xu, Wenjia, et autres
Publié: (2024)
par: Xu, Wenjia, et autres
Publié: (2024)
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
par: Wu, Diankun, et autres
Publié: (2025)
par: Wu, Diankun, et autres
Publié: (2025)
Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
par: Song, Yeji, et autres
Publié: (2024)
par: Song, Yeji, et autres
Publié: (2024)
AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
par: Ahn, Sunghyun, et autres
Publié: (2025)
par: Ahn, Sunghyun, et autres
Publié: (2025)
Memory-based Adapters for Online 3D Scene Perception
par: Xu, Xiuwei, et autres
Publié: (2024)
par: Xu, Xiuwei, et autres
Publié: (2024)
Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model
par: Zhang, Shengjun, et autres
Publié: (2025)
par: Zhang, Shengjun, et autres
Publié: (2025)
Zero-Shot Scene Change Detection
par: Cho, Kyusik, et autres
Publié: (2024)
par: Cho, Kyusik, et autres
Publié: (2024)
EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning
par: Zhang, Xiao, et autres
Publié: (2025)
par: Zhang, Xiao, et autres
Publié: (2025)
SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
par: Chen, Weiliang, et autres
Publié: (2025)
par: Chen, Weiliang, et autres
Publié: (2025)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
par: Li, Xiaolou, et autres
Publié: (2024)
par: Li, Xiaolou, et autres
Publié: (2024)
Semantic Flow: Learning Semantic Field of Dynamic Scenes from Monocular Videos
par: Tian, Fengrui, et autres
Publié: (2024)
par: Tian, Fengrui, et autres
Publié: (2024)
SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation
par: Zhang, Jiwen, et autres
Publié: (2026)
par: Zhang, Jiwen, et autres
Publié: (2026)
Zero-Shot Hashing Based on Reconstruction With Part Alignment
par: Jiang, Yan, et autres
Publié: (2025)
par: Jiang, Yan, et autres
Publié: (2025)
OnlineAnySeg: Online Zero-Shot 3D Segmentation by Visual Foundation Model Guided 2D Mask Merging
par: Tang, Yijie, et autres
Publié: (2025)
par: Tang, Yijie, et autres
Publié: (2025)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
par: Zhang, Jiwen, et autres
Publié: (2026)
par: Zhang, Jiwen, et autres
Publié: (2026)
SAMJAM: Zero-Shot Video Scene Graph Generation for Egocentric Kitchen Videos
par: Li, Joshua, et autres
Publié: (2025)
par: Li, Joshua, et autres
Publié: (2025)
Learning by Imagining: Debiased Feature Augmentation for Compositional Zero-Shot Learning
par: Zhang, Haozhe, et autres
Publié: (2025)
par: Zhang, Haozhe, et autres
Publié: (2025)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
par: Gkanatsios, Nikolaos, et autres
Publié: (2023)
par: Gkanatsios, Nikolaos, et autres
Publié: (2023)
Revisiting 3D Reconstruction Kernels as Low-Pass Filters
par: Zhang, Shengjun, et autres
Publié: (2026)
par: Zhang, Shengjun, et autres
Publié: (2026)
Anything in Any Scene: Photorealistic Video Object Insertion
par: Bai, Chen, et autres
Publié: (2024)
par: Bai, Chen, et autres
Publié: (2024)
ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation
par: Li, Hongjie, et autres
Publié: (2024)
par: Li, Hongjie, et autres
Publié: (2024)
Zero-Shot Personalization of Objects via Textual Inversion
par: Roy, Aniket, et autres
Publié: (2026)
par: Roy, Aniket, et autres
Publié: (2026)
Zero-Shot Temporal Interaction Localization for Egocentric Videos
par: Zhang, Erhang, et autres
Publié: (2025)
par: Zhang, Erhang, et autres
Publié: (2025)
Learning Visual Proxy for Compositional Zero-Shot Learning
par: Zhang, Shiyu, et autres
Publié: (2025)
par: Zhang, Shiyu, et autres
Publié: (2025)
Visual Adaptive Prompting for Compositional Zero-Shot Learning
par: Stein, Kyle, et autres
Publié: (2025)
par: Stein, Kyle, et autres
Publié: (2025)
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
par: Liu, Fangfu, et autres
Publié: (2025)
par: Liu, Fangfu, et autres
Publié: (2025)
Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera
par: Guo, Yuliang, et autres
Publié: (2025)
par: Guo, Yuliang, et autres
Publié: (2025)
EmboAlign: Aligning Video Generation with Compositional Constraints for Zero-Shot Manipulation
par: Zhang, Gehao, et autres
Publié: (2026)
par: Zhang, Gehao, et autres
Publié: (2026)
RARE: Refine Any Registration of Pairwise Point Clouds via Zero-Shot Learning
par: Zheng, Chengyu, et autres
Publié: (2025)
par: Zheng, Chengyu, et autres
Publié: (2025)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
par: Yang, Shuai, et autres
Publié: (2024)
par: Yang, Shuai, et autres
Publié: (2024)
OnlineX: Unified Online 3D Reconstruction and Understanding with Active-to-Stable State Evolution
par: Xia, Chong, et autres
Publié: (2026)
par: Xia, Chong, et autres
Publié: (2026)
TextPSG: Panoptic Scene Graph Generation from Textual Descriptions
par: Zhao, Chengyang, et autres
Publié: (2023)
par: Zhao, Chengyang, et autres
Publié: (2023)
Coherent Zero-Shot Visual Instruction Generation
par: Phung, Quynh, et autres
Publié: (2024)
par: Phung, Quynh, et autres
Publié: (2024)
Zero-Shot Temporal Action Localization Through Textual Guidance
par: Liberatori, Benedetta, et autres
Publié: (2026)
par: Liberatori, Benedetta, et autres
Publié: (2026)
Documents similaires
-
SimRecon: SimReady Compositional Scene Reconstruction from Real Videos
par: Xia, Chong, et autres
Publié: (2026) -
ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment
par: Xia, Chong, et autres
Publié: (2025) -
ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
par: Liu, Fangfu, et autres
Publié: (2024) -
VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
par: Wang, Hanyang, et autres
Publié: (2025) -
GranAlign: Granularity-Aware Alignment Framework for Zero-Shot Video Moment Retrieval
par: Jeon, Mingyu, et autres
Publié: (2026)