Large Spatial Model: End-to-end Unposed Images to Semantic 3D
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Zhiwen, Zhang, Jian, Cong, Wenyan, Wang, Peihao, Li, Renjie, Wen, Kairun, Zhou, Shijie, Kadambi, Achuta, Wang, Zhangyang, Xu, Danfei, Ivanovic, Boris, Pavone, Marco, Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InstantSplat: Sparse-view Gaussian Splatting in Seconds
by: Fan, Zhiwen, et al.
Published: (2024)
by: Fan, Zhiwen, et al.
Published: (2024)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
by: Cong, Wenyan, et al.
Published: (2025)
by: Cong, Wenyan, et al.
Published: (2025)
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
4K4DGen: Panoramic 4D Generation at 4K Resolution
by: Li, Renjie, et al.
Published: (2024)
by: Li, Renjie, et al.
Published: (2024)
LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPS
by: Fan, Zhiwen, et al.
Published: (2023)
by: Fan, Zhiwen, et al.
Published: (2023)
Learning Traffic Crashes as Language: Datasets, Benchmarks, and What-if Causal Analyses
by: Fan, Zhiwen, et al.
Published: (2024)
by: Fan, Zhiwen, et al.
Published: (2024)
LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models
by: Lu, Ziqi, et al.
Published: (2024)
by: Lu, Ziqi, et al.
Published: (2024)
Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
by: Zhou, Shijie, et al.
Published: (2023)
by: Zhou, Shijie, et al.
Published: (2023)
Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving
by: Ivanovic, Boris, et al.
Published: (2025)
by: Ivanovic, Boris, et al.
Published: (2025)
Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning
by: Wang, Peihao, et al.
Published: (2025)
by: Wang, Peihao, et al.
Published: (2025)
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
by: Yang, Jiawei, et al.
Published: (2025)
by: Yang, Jiawei, et al.
Published: (2025)
Can Test-Time Scaling Improve World Foundation Model?
by: Cong, Wenyan, et al.
Published: (2025)
by: Cong, Wenyan, et al.
Published: (2025)
DreamDrive: Generative 4D Scene Modeling from Street View Images
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
by: Jeon, Changwoo, et al.
Published: (2026)
by: Jeon, Changwoo, et al.
Published: (2026)
Lift3D: Zero-Shot Lifting of Any 2D Vision Model to 3D
by: T, Mukund Varma, et al.
Published: (2024)
by: T, Mukund Varma, et al.
Published: (2024)
Position: Weight Space Should Be a First-Class Generative AI Modality
by: Wang, Zhangyang, et al.
Published: (2026)
by: Wang, Zhangyang, et al.
Published: (2026)
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
by: Li, Longfei, et al.
Published: (2025)
by: Li, Longfei, et al.
Published: (2025)
The Potential and Perils of Generative Artificial Intelligence for Quality Improvement and Patient Safety
by: Jalilian, Laleh, et al.
Published: (2024)
by: Jalilian, Laleh, et al.
Published: (2024)
Generalization Error Analysis for Sparse Mixture-of-Experts: A Preliminary Study
by: Zhao, Jinze, et al.
Published: (2024)
by: Zhao, Jinze, et al.
Published: (2024)
InstantRestore: Single-Step Personalized Face Restoration with Shared-Image Attention
by: Zhang, Howard, et al.
Published: (2024)
by: Zhang, Howard, et al.
Published: (2024)
MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
by: He, Xuehai, et al.
Published: (2025)
by: He, Xuehai, et al.
Published: (2025)
Parallelized Spatiotemporal Binding
by: Singh, Gautam, et al.
Published: (2024)
by: Singh, Gautam, et al.
Published: (2024)
Driving Everywhere with Large Language Model Policy Adaptation
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
Solutions to Deepfakes: Can Camera Hardware, Cryptography, and Deep Learning Verify Real Images?
by: Vilesov, Alexander, et al.
Published: (2024)
by: Vilesov, Alexander, et al.
Published: (2024)
Latent Chain-of-Thought World Modeling for End-to-End Driving
by: Tan, Shuhan, et al.
Published: (2025)
by: Tan, Shuhan, et al.
Published: (2025)
Producing and Leveraging Online Map Uncertainty in Trajectory Prediction
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Accelerating Online Mapping and Behavior Prediction via Direct BEV Feature Attention
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Data Scaling Laws for End-to-End Autonomous Driving
by: Naumann, Alexander, et al.
Published: (2025)
by: Naumann, Alexander, et al.
Published: (2025)
STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes
by: Yang, Jiawei, et al.
Published: (2024)
by: Yang, Jiawei, et al.
Published: (2024)
VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment
by: Cong, Wenyan, et al.
Published: (2025)
by: Cong, Wenyan, et al.
Published: (2025)
Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving
by: Tian, Ran, et al.
Published: (2024)
by: Tian, Ran, et al.
Published: (2024)
SparseGS: Real-Time 360° Sparse View Synthesis using Gaussian Splatting
by: Xiong, Haolin, et al.
Published: (2023)
by: Xiong, Haolin, et al.
Published: (2023)
FSGS: Real-Time Few-shot View Synthesis using Gaussian Splatting
by: Zhu, Zehao, et al.
Published: (2023)
by: Zhu, Zehao, et al.
Published: (2023)
Oscillation Inversion: Understand the structure of Large Flow Model through the Lens of Inversion Method
by: Zheng, Yan, et al.
Published: (2024)
by: Zheng, Yan, et al.
Published: (2024)
DTPP: Differentiable Joint Conditional Prediction and Cost Evaluation for Tree Policy Planning in Autonomous Driving
by: Huang, Zhiyu, et al.
Published: (2023)
by: Huang, Zhiyu, et al.
Published: (2023)
SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images
by: Sheng, Yu, et al.
Published: (2025)
by: Sheng, Yu, et al.
Published: (2025)
Equitable non-contact infrared thermography after solar loading using deep learning
by: Zhao, Ellin Q., et al.
Published: (2023)
by: Zhao, Ellin Q., et al.
Published: (2023)
Meta ControlNet: Enhancing Task Adaptation via Meta Learning
by: Yang, Junjie, et al.
Published: (2023)
by: Yang, Junjie, et al.
Published: (2023)
Similar Items
-
InstantSplat: Sparse-view Gaussian Splatting in Seconds
by: Fan, Zhiwen, et al.
Published: (2024) -
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026) -
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
by: Cong, Wenyan, et al.
Published: (2025) -
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
by: Zhou, Shijie, et al.
Published: (2024) -
4K4DGen: Panoramic 4D Generation at 4K Resolution
by: Li, Renjie, et al.
Published: (2024)