Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, David Yifan, Zhai, Albert J., Wang, Shenlong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
by: Liu, Shaowei, et al.
Published: (2025)
by: Liu, Shaowei, et al.
Published: (2025)
UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
by: Su, Zhaolong, et al.
Published: (2025)
by: Su, Zhaolong, et al.
Published: (2025)
Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE
by: Zhu, Xun, et al.
Published: (2024)
by: Zhu, Xun, et al.
Published: (2024)
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
by: Liu, Shaowei, et al.
Published: (2024)
by: Liu, Shaowei, et al.
Published: (2024)
Stag-1: Towards Realistic 4D Driving Simulation with Video Generation Model
by: Wang, Lening, et al.
Published: (2024)
by: Wang, Lening, et al.
Published: (2024)
UniFusion: Vision-Language Model as Unified Encoder in Image Generation
by: Li, Kevin, et al.
Published: (2025)
by: Li, Kevin, et al.
Published: (2025)
Any4D: Unified Feed-Forward Metric 4D Reconstruction
by: Karhade, Jay, et al.
Published: (2025)
by: Karhade, Jay, et al.
Published: (2025)
Streaming 4D Visual Geometry Transformer
by: Zhuo, Dong, et al.
Published: (2025)
by: Zhuo, Dong, et al.
Published: (2025)
3D-LFM: Lifting Foundation Model
by: Dabhi, Mosam, et al.
Published: (2023)
by: Dabhi, Mosam, et al.
Published: (2023)
Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion
by: Mou, Linzhan, et al.
Published: (2024)
by: Mou, Linzhan, et al.
Published: (2024)
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
Pixels to Play: A Foundation Model for 3D Gameplay
by: Yue, Yuguang, et al.
Published: (2025)
by: Yue, Yuguang, et al.
Published: (2025)
Training Video Foundation Models with NVIDIA NeMo
by: Patel, Zeeshan, et al.
Published: (2025)
by: Patel, Zeeshan, et al.
Published: (2025)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
An Investigation of Visual Foundation Models Robustness
by: Gupta, Sandeep, et al.
Published: (2025)
by: Gupta, Sandeep, et al.
Published: (2025)
VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
by: Barreto, Jesimon, et al.
Published: (2025)
by: Barreto, Jesimon, et al.
Published: (2025)
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
by: Chen, Leon Liangyu, et al.
Published: (2026)
by: Chen, Leon Liangyu, et al.
Published: (2026)
Revisiting Model Stitching In the Foundation Model Era
by: Mai, Zheda, et al.
Published: (2026)
by: Mai, Zheda, et al.
Published: (2026)
Physical Property Understanding from Language-Embedded Feature Fields
by: Zhai, Albert J., et al.
Published: (2024)
by: Zhai, Albert J., et al.
Published: (2024)
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
by: Li, Edward, et al.
Published: (2025)
by: Li, Edward, et al.
Published: (2025)
Geometric 4D Stitching for Grounded 4D Generation
by: Park, Sunwoo, et al.
Published: (2026)
by: Park, Sunwoo, et al.
Published: (2026)
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
by: Mai, Zheda, et al.
Published: (2025)
by: Mai, Zheda, et al.
Published: (2025)
Scaling 4D Representations
by: Carreira, João, et al.
Published: (2024)
by: Carreira, João, et al.
Published: (2024)
Unified Supervision For Vision-Language Modeling in 3D Computed Tomography
by: Lee, Hao-Chih, et al.
Published: (2025)
by: Lee, Hao-Chih, et al.
Published: (2025)
Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
by: Zhai, Guangyao, et al.
Published: (2025)
by: Zhai, Guangyao, et al.
Published: (2025)
World Simulation with Video Foundation Models for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
Geometry-aware 4D Video Generation for Robot Manipulation
by: Liu, Zeyi, et al.
Published: (2025)
by: Liu, Zeyi, et al.
Published: (2025)
ODIN: A Single Model for 2D and 3D Segmentation
by: Jain, Ayush, et al.
Published: (2024)
by: Jain, Ayush, et al.
Published: (2024)
Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
by: Yang, Zhijian, et al.
Published: (2025)
by: Yang, Zhijian, et al.
Published: (2025)
PhiNet v2: A Mask-Free Brain-Inspired Vision Foundation Model from Video
by: Yamada, Makoto, et al.
Published: (2025)
by: Yamada, Makoto, et al.
Published: (2025)
SuperGaussian: Repurposing Video Models for 3D Super Resolution
by: Shen, Yuan, et al.
Published: (2024)
by: Shen, Yuan, et al.
Published: (2024)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
by: Jeong, Hyeonho, et al.
Published: (2024)
by: Jeong, Hyeonho, et al.
Published: (2024)
Unified Lexical Representation for Interpretable Visual-Language Alignment
by: Li, Yifan, et al.
Published: (2024)
by: Li, Yifan, et al.
Published: (2024)
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
by: Arkhipkin, Vladimir, et al.
Published: (2025)
by: Arkhipkin, Vladimir, et al.
Published: (2025)
Research on the Spatial Data Intelligent Foundation Model
by: Wang, Shaohua, et al.
Published: (2024)
by: Wang, Shaohua, et al.
Published: (2024)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
by: Liu, Yunong, et al.
Published: (2024)
by: Liu, Yunong, et al.
Published: (2024)
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
by: Peng, Shangpin, et al.
Published: (2025)
by: Peng, Shangpin, et al.
Published: (2025)
LRM: Large Reconstruction Model for Single Image to 3D
by: Hong, Yicong, et al.
Published: (2023)
by: Hong, Yicong, et al.
Published: (2023)
Similar Items
-
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
by: Liu, Shaowei, et al.
Published: (2025) -
UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
by: Su, Zhaolong, et al.
Published: (2025) -
Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE
by: Zhu, Xun, et al.
Published: (2024) -
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
by: Liu, Shaowei, et al.
Published: (2024) -
Stag-1: Towards Realistic 4D Driving Simulation with Video Generation Model
by: Wang, Lening, et al.
Published: (2024)