UVRM: A Scalable 3D Reconstruction Model from Unposed Videos
Fuente:
arXiv
Guardado en:
| Autores principales: | Kao, Shiu-hong, Li, Xiao, Wang, Jinglu, Li, Yang, Tang, Chi-Keung, Tai, Yu-Wing, Lu, Yan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
por: Kao, Shiu-hong, et al.
Publicado: (2025)
por: Kao, Shiu-hong, et al.
Publicado: (2025)
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
por: Kao, Shiu-hong, et al.
Publicado: (2025)
por: Kao, Shiu-hong, et al.
Publicado: (2025)
InceptionHuman: Controllable Prompt-to-NeRF for Photorealistic 3D Human Generation
por: Kao, Shiu-hong, et al.
Publicado: (2023)
por: Kao, Shiu-hong, et al.
Publicado: (2023)
StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image Streams
por: LI, Yang, et al.
Publicado: (2025)
por: LI, Yang, et al.
Publicado: (2025)
Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-Observations for High-Quality Sparse-View Reconstruction
por: Liu, Xinhang, et al.
Publicado: (2023)
por: Liu, Xinhang, et al.
Publicado: (2023)
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
por: Kao, Shiu-hong, et al.
Publicado: (2026)
por: Kao, Shiu-hong, et al.
Publicado: (2026)
FED-NeRF: Achieve High 3D Consistency and Temporal Coherence for Face Video Editing on Dynamic NeRF
por: Zhang, Hao, et al.
Publicado: (2024)
por: Zhang, Hao, et al.
Publicado: (2024)
Multimodal Generation of Animatable 3D Human Models with AvatarForge
por: Liu, Xinhang, et al.
Publicado: (2025)
por: Liu, Xinhang, et al.
Publicado: (2025)
Agentic 3D Scene Generation with Spatially Contextualized VLMs
por: Liu, Xinhang, et al.
Publicado: (2025)
por: Liu, Xinhang, et al.
Publicado: (2025)
WorldCraft: Photo-Realistic 3D World Creation and Customization via LLM Agents
por: Liu, Xinhang, et al.
Publicado: (2025)
por: Liu, Xinhang, et al.
Publicado: (2025)
ChatCam: Empowering Camera Control through Conversational AI
por: Liu, Xinhang, et al.
Publicado: (2024)
por: Liu, Xinhang, et al.
Publicado: (2024)
ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation
por: Wang, Zixuan, et al.
Publicado: (2025)
por: Wang, Zixuan, et al.
Publicado: (2025)
DragVideo: Interactive Drag-style Video Editing
por: Deng, Yufan, et al.
Publicado: (2023)
por: Deng, Yufan, et al.
Publicado: (2023)
Inpaint4DNeRF: Promptable Spatio-Temporal NeRF Inpainting with Generative Diffusion Models
por: Jiang, Han, et al.
Publicado: (2023)
por: Jiang, Han, et al.
Publicado: (2023)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
por: Liu, Xinhang, et al.
Publicado: (2025)
por: Liu, Xinhang, et al.
Publicado: (2025)
SANeRF-HQ: Segment Anything for NeRF in High Quality
por: Liu, Yichen, et al.
Publicado: (2023)
por: Liu, Yichen, et al.
Publicado: (2023)
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
por: Wang, Zixuan, et al.
Publicado: (2024)
por: Wang, Zixuan, et al.
Publicado: (2024)
VP-LLM: Text-Driven 3D Volume Completion with Large Language Models through Patchification
por: Liu, Jianmeng, et al.
Publicado: (2024)
por: Liu, Jianmeng, et al.
Publicado: (2024)
ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation
por: Ao, Yuzhuo, et al.
Publicado: (2026)
por: Ao, Yuzhuo, et al.
Publicado: (2026)
Beyond and Free from Diffusion: Invertible Guided Consistency Training
por: Hsu, Chia-Hong, et al.
Publicado: (2025)
por: Hsu, Chia-Hong, et al.
Publicado: (2025)
StableKD: Breaking Inter-block Optimization Entanglement for Stable Knowledge Distillation
por: Kao, Shiu-hong, et al.
Publicado: (2023)
por: Kao, Shiu-hong, et al.
Publicado: (2023)
GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
por: Zhou, Junwei, et al.
Publicado: (2025)
por: Zhou, Junwei, et al.
Publicado: (2025)
Unposed 3DGS Reconstruction with Probabilistic Procrustes Mapping
por: Cheng, Chong, et al.
Publicado: (2025)
por: Cheng, Chong, et al.
Publicado: (2025)
LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
por: Lin, Chin-Yang, et al.
Publicado: (2025)
por: Lin, Chin-Yang, et al.
Publicado: (2025)
Navigating Motion Agents in Dynamic and Cluttered Environments through LLM Reasoning
por: Zhao, Yubo, et al.
Publicado: (2025)
por: Zhao, Yubo, et al.
Publicado: (2025)
UniSem: Generalizable Semantic 3D Reconstruction from Sparse Unposed Images
por: Liao, Guibiao, et al.
Publicado: (2026)
por: Liao, Guibiao, et al.
Publicado: (2026)
Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs
por: Wu, Qi, et al.
Publicado: (2024)
por: Wu, Qi, et al.
Publicado: (2024)
LucidFusion: Reconstructing 3D Gaussians with Arbitrary Unposed Images
por: He, Hao, et al.
Publicado: (2024)
por: He, Hao, et al.
Publicado: (2024)
Distill Gold from Massive Ores: Bi-level Data Pruning towards Efficient Dataset Distillation
por: Xu, Yue, et al.
Publicado: (2023)
por: Xu, Yue, et al.
Publicado: (2023)
SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
por: Huang-Menders, Alexander, et al.
Publicado: (2025)
por: Huang-Menders, Alexander, et al.
Publicado: (2025)
GS-Marker: Generalizable and Robust Watermarking for 3D Gaussian Splatting
por: Li, Lijiang, et al.
Publicado: (2025)
por: Li, Lijiang, et al.
Publicado: (2025)
DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images
por: Chen, Xiaoxue, et al.
Publicado: (2025)
por: Chen, Xiaoxue, et al.
Publicado: (2025)
Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation
por: Zhou, Junwei, et al.
Publicado: (2026)
por: Zhou, Junwei, et al.
Publicado: (2026)
TransmissiveGS: Residual-Guided Disentangled Gaussian Splatting for Transmissive Scene Reconstruction and Rendering
por: Liang, Zhenyu, et al.
Publicado: (2026)
por: Liang, Zhenyu, et al.
Publicado: (2026)
KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos
por: Chou, Gene, et al.
Publicado: (2024)
por: Chou, Gene, et al.
Publicado: (2024)
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
por: Fan, Zhiwen, et al.
Publicado: (2024)
por: Fan, Zhiwen, et al.
Publicado: (2024)
Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
por: Liu, Jiageng, et al.
Publicado: (2025)
por: Liu, Jiageng, et al.
Publicado: (2025)
Pragmatist: Multiview Conditional Diffusion Models for High-Fidelity 3D Reconstruction from Unposed Sparse Views
por: Zhang, Songchun, et al.
Publicado: (2024)
por: Zhang, Songchun, et al.
Publicado: (2024)
Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sampling
por: Liu, Xinhang, et al.
Publicado: (2024)
por: Liu, Xinhang, et al.
Publicado: (2024)
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
por: Hur, Junhwa, et al.
Publicado: (2026)
por: Hur, Junhwa, et al.
Publicado: (2026)
Ejemplares similares
-
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
por: Kao, Shiu-hong, et al.
Publicado: (2025) -
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
por: Kao, Shiu-hong, et al.
Publicado: (2025) -
InceptionHuman: Controllable Prompt-to-NeRF for Photorealistic 3D Human Generation
por: Kao, Shiu-hong, et al.
Publicado: (2023) -
StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image Streams
por: LI, Yang, et al.
Publicado: (2025) -
Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-Observations for High-Quality Sparse-View Reconstruction
por: Liu, Xinhang, et al.
Publicado: (2023)