NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yuxue, Fan, Lue, Shi, Ziqi, Peng, Junran, Wang, Feng, Zhang, Zhaoxiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MixSup: Mixed-grained Supervision for Label-efficient LiDAR-based 3D Object Detection
by: Yang, Yuxue, et al.
Published: (2024)
by: Yang, Yuxue, et al.
Published: (2024)
LayerAnimate: Layer-level Control for Animation
by: Yang, Yuxue, et al.
Published: (2025)
by: Yang, Yuxue, et al.
Published: (2025)
Trim 3D Gaussian Splatting for Accurate Geometry Representation
by: Fan, Lue, et al.
Published: (2024)
by: Fan, Lue, et al.
Published: (2024)
SAGD: Boundary-Enhanced Segment Anything in 3D Gaussian via Gaussian Decomposition
by: Hu, Xu, et al.
Published: (2024)
by: Hu, Xu, et al.
Published: (2024)
TC-Light: Temporally Coherent Generative Rendering for Realistic World Transfer
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
CityGaussian: Real-time High-quality Large-Scale Scene Rendering with Gaussians
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Enhancing End-to-End Autonomous Driving with Latent World Model
by: Li, Yingyan, et al.
Published: (2024)
by: Li, Yingyan, et al.
Published: (2024)
EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models
by: Zhang, Yixuan, et al.
Published: (2025)
by: Zhang, Yixuan, et al.
Published: (2025)
End-to-End Driving with Online Trajectory Evaluation via BEV World Model
by: Li, Yingyan, et al.
Published: (2025)
by: Li, Yingyan, et al.
Published: (2025)
DeepVerse: 4D Autoregressive Video Generation as a World Model
by: Chen, Junyi, et al.
Published: (2025)
by: Chen, Junyi, et al.
Published: (2025)
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
by: Zheng, Sixiao, et al.
Published: (2026)
by: Zheng, Sixiao, et al.
Published: (2026)
OOD-HOI: Text-Driven 3D Whole-Body Human-Object Interactions Generation Beyond Training Domains
by: Zhang, Yixuan, et al.
Published: (2024)
by: Zhang, Yixuan, et al.
Published: (2024)
CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
VGGT-X: When VGGT Meets Dense Novel View Synthesis
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object Detection
by: Zhang, Guowen, et al.
Published: (2024)
by: Zhang, Guowen, et al.
Published: (2024)
DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
by: Wen, Kairun, et al.
Published: (2025)
by: Wen, Kairun, et al.
Published: (2025)
Fully Sparse Fusion for 3D Object Detection
by: Li, Yingyan, et al.
Published: (2023)
by: Li, Yingyan, et al.
Published: (2023)
FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes
by: Fan, Lue, et al.
Published: (2024)
by: Fan, Lue, et al.
Published: (2024)
GFlow: Recovering 4D World from Monocular Video
by: Wang, Shizun, et al.
Published: (2024)
by: Wang, Shizun, et al.
Published: (2024)
FreeVS: Generative View Synthesis on Free Driving Trajectory
by: Wang, Qitai, et al.
Published: (2024)
by: Wang, Qitai, et al.
Published: (2024)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
GSPlane: Concise and Accurate Planar Reconstruction via Structured Representation
by: Gan, Ruitong, et al.
Published: (2025)
by: Gan, Ruitong, et al.
Published: (2025)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
DomainVerse: A Benchmark Towards Real-World Distribution Shifts For Tuning-Free Adaptive Domain Generalization
by: Hou, Feng, et al.
Published: (2024)
by: Hou, Feng, et al.
Published: (2024)
WonderVerse: Extendable 3D Scene Generation with Video Generative Models
by: Feng, Hao, et al.
Published: (2025)
by: Feng, Hao, et al.
Published: (2025)
Predicting 4D Hand Trajectory from Monocular Videos
by: Ye, Yufei, et al.
Published: (2025)
by: Ye, Yufei, et al.
Published: (2025)
FlexDrive: Toward Trajectory Flexibility in Driving Scene Reconstruction and Rendering
by: Zhou, Jingqiu, et al.
Published: (2025)
by: Zhou, Jingqiu, et al.
Published: (2025)
Cross360: 360° Monocular Depth Estimation via Cross Projections Across Scales
by: Huang, Kun, et al.
Published: (2026)
by: Huang, Kun, et al.
Published: (2026)
Seek for Incantations: Towards Accurate Text-to-Image Diffusion Synthesis through Prompt Engineering
by: Yu, Chang, et al.
Published: (2024)
by: Yu, Chang, et al.
Published: (2024)
UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
WorldTree: Towards 4D Dynamic Worlds from Monocular Video using Tree-Chains
by: Wang, Qisen, et al.
Published: (2026)
by: Wang, Qisen, et al.
Published: (2026)
Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
by: Huang, Jiaxin, et al.
Published: (2025)
by: Huang, Jiaxin, et al.
Published: (2025)
FurniScene: A Large-scale 3D Room Dataset with Intricate Furnishing Scenes
by: Zhang, Genghao, et al.
Published: (2024)
by: Zhang, Genghao, et al.
Published: (2024)
Monocular Occupancy Prediction for Scalable Indoor Scenes
by: Yu, Hongxiao, et al.
Published: (2024)
by: Yu, Hongxiao, et al.
Published: (2024)
4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos
by: Xu, Zhen, et al.
Published: (2025)
by: Xu, Zhen, et al.
Published: (2025)
GA-Drive: Geometry-Appearance Decoupled Modeling for Free-viewpoint Driving Scene Generation
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
MCOP: Multi-UAV Collaborative Occupancy Prediction
by: Lin, Zefu, et al.
Published: (2025)
by: Lin, Zefu, et al.
Published: (2025)
ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World
by: Dhiman, Ankit, et al.
Published: (2025)
by: Dhiman, Ankit, et al.
Published: (2025)
Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos
by: Luo, Xianrui, et al.
Published: (2025)
by: Luo, Xianrui, et al.
Published: (2025)
Similar Items
-
MixSup: Mixed-grained Supervision for Label-efficient LiDAR-based 3D Object Detection
by: Yang, Yuxue, et al.
Published: (2024) -
LayerAnimate: Layer-level Control for Animation
by: Yang, Yuxue, et al.
Published: (2025) -
Trim 3D Gaussian Splatting for Accurate Geometry Representation
by: Fan, Lue, et al.
Published: (2024) -
SAGD: Boundary-Enhanced Segment Anything in 3D Gaussian via Gaussian Decomposition
by: Hu, Xu, et al.
Published: (2024) -
TC-Light: Temporally Coherent Generative Rendering for Realistic World Transfer
by: Liu, Yang, et al.
Published: (2025)