E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Qitao, Tan, Hao, Wang, Qianqian, Bi, Sai, Zhang, Kai, Sunkavalli, Kalyan, Tulsiani, Shubham, Jiang, Hanwen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RayZer: A Self-supervised Large View Synthesis Model
by: Jiang, Hanwen, et al.
Published: (2025)
by: Jiang, Hanwen, et al.
Published: (2025)
Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis
by: Zhao, Qitao, et al.
Published: (2024)
by: Zhao, Qitao, et al.
Published: (2024)
GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting
by: Zhang, Kai, et al.
Published: (2024)
by: Zhang, Kai, et al.
Published: (2024)
Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning
by: Cong, Zhongxiao, et al.
Published: (2026)
by: Cong, Zhongxiao, et al.
Published: (2026)
DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion
by: Zhao, Qitao, et al.
Published: (2025)
by: Zhao, Qitao, et al.
Published: (2025)
MeshLRM: Large Reconstruction Model for High-Quality Meshes
by: Wei, Xinyue, et al.
Published: (2024)
by: Wei, Xinyue, et al.
Published: (2024)
LRM: Large Reconstruction Model for Single Image to 3D
by: Hong, Yicong, et al.
Published: (2023)
by: Hong, Yicong, et al.
Published: (2023)
tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction
by: Wang, Chen, et al.
Published: (2026)
by: Wang, Chen, et al.
Published: (2026)
WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments
by: Chen, Xuweiyi, et al.
Published: (2026)
by: Chen, Xuweiyi, et al.
Published: (2026)
NeuManifold: Neural Watertight Manifold Reconstruction with Efficient and High-Quality Rendering Support
by: Wei, Xinyue, et al.
Published: (2023)
by: Wei, Xinyue, et al.
Published: (2023)
Turbo3D: Ultra-fast Text-to-3D Generation
by: Hu, Hanzhe, et al.
Published: (2024)
by: Hu, Hanzhe, et al.
Published: (2024)
Test-Time Training Done Right
by: Zhang, Tianyuan, et al.
Published: (2025)
by: Zhang, Tianyuan, et al.
Published: (2025)
4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
by: Ma, Ziqiao, et al.
Published: (2025)
by: Ma, Ziqiao, et al.
Published: (2025)
DressRecon: Freeform 4D Human Reconstruction from Monocular Video
by: Tan, Jeff, et al.
Published: (2024)
by: Tan, Jeff, et al.
Published: (2024)
DemoDiffusion: One-Shot Human Imitation using pre-trained Diffusion Policy
by: Park, Sungjae, et al.
Published: (2025)
by: Park, Sungjae, et al.
Published: (2025)
Neural Directional Encoding for Efficient and Accurate View-Dependent Appearance Modeling
by: Wu, Liwen, et al.
Published: (2024)
by: Wu, Liwen, et al.
Published: (2024)
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
by: Wu, Yu, et al.
Published: (2026)
by: Wu, Yu, et al.
Published: (2026)
Zer
Published: (2022)
Published: (2022)
G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis
by: Ye, Yufei, et al.
Published: (2024)
by: Ye, Yufei, et al.
Published: (2024)
Self-supervised Pre-training of Text Recognizers
by: Kišš, Martin, et al.
Published: (2024)
by: Kišš, Martin, et al.
Published: (2024)
Cameras as Rays: Pose Estimation via Ray Diffusion
by: Zhang, Jason Y., et al.
Published: (2024)
by: Zhang, Jason Y., et al.
Published: (2024)
SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generation
by: Bokhovkin, Alexey, et al.
Published: (2024)
by: Bokhovkin, Alexey, et al.
Published: (2024)
LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias
by: Jin, Haian, et al.
Published: (2024)
by: Jin, Haian, et al.
Published: (2024)
Transferable Watermarking to Self-supervised Pre-trained Graph Encoders by Trigger Embeddings
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
by: Vuong, Khiem, et al.
Published: (2025)
by: Vuong, Khiem, et al.
Published: (2025)
Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation
by: Kuang, Yuxuan, et al.
Published: (2026)
by: Kuang, Yuxuan, et al.
Published: (2026)
MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation
by: Hu, Hanzhe, et al.
Published: (2024)
by: Hu, Hanzhe, et al.
Published: (2024)
MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data
by: Jiang, Hanwen, et al.
Published: (2024)
by: Jiang, Hanwen, et al.
Published: (2024)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
by: Peng, Junyi, et al.
Published: (2024)
by: Peng, Junyi, et al.
Published: (2024)
Improving Self-supervised Pre-training using Accent-Specific Codebooks
by: Prabhu, Darshan, et al.
Published: (2024)
by: Prabhu, Darshan, et al.
Published: (2024)
LightIt: Illumination Modeling and Control for Diffusion Models
by: Kocsis, Peter, et al.
Published: (2024)
by: Kocsis, Peter, et al.
Published: (2024)
Foundation Model for Endoscopy Video Analysis via Large-scale Self-supervised Pre-train
by: Wang, Zhao, et al.
Published: (2023)
by: Wang, Zhao, et al.
Published: (2023)
Multi-modal Cross-domain Self-supervised Pre-training for fMRI and EEG Fusion
by: Wei, Xinxu, et al.
Published: (2024)
by: Wei, Xinxu, et al.
Published: (2024)
GraphSculptor: Sculpting Pre-training Coreset for Graph Self-supervised Learning
by: Liu, Chuang, et al.
Published: (2026)
by: Liu, Chuang, et al.
Published: (2026)
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
by: Bawazir, Ameera, et al.
Published: (2024)
by: Bawazir, Ameera, et al.
Published: (2024)
Pre-trained Vision-Language Models Learn Discoverable Visual Concepts
by: Zang, Yuan, et al.
Published: (2024)
by: Zang, Yuan, et al.
Published: (2024)
Diophantine $D(n)$-quadruples in $\mathbb{Z}[\sqrt{4k + 2}]$
by: Chakraborty, Kalyan, et al.
Published: (2023)
by: Chakraborty, Kalyan, et al.
Published: (2023)
SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-supervised Learning
by: Lv, Peizhuo, et al.
Published: (2022)
by: Lv, Peizhuo, et al.
Published: (2022)
Structure-aware World Model for Probe Guidance via Large-scale Self-supervised Pre-train
by: Jiang, Haojun, et al.
Published: (2024)
by: Jiang, Haojun, et al.
Published: (2024)
LightSwitch: Multi-view Relighting with Material-guided Diffusion
by: Litman, Yehonathan, et al.
Published: (2025)
by: Litman, Yehonathan, et al.
Published: (2025)
Similar Items
-
RayZer: A Self-supervised Large View Synthesis Model
by: Jiang, Hanwen, et al.
Published: (2025) -
Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis
by: Zhao, Qitao, et al.
Published: (2024) -
GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting
by: Zhang, Kai, et al.
Published: (2024) -
Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning
by: Cong, Zhongxiao, et al.
Published: (2026) -
DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion
by: Zhao, Qitao, et al.
Published: (2025)