Streaming 4D Visual Geometry Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Zhuo, Dong, Zheng, Wenzhao, Guo, Jiahe, Wu, Yuqi, Zhou, Jie, Lu, Jiwen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
by: Wu, Yuqi, et al.
Published: (2025)
by: Wu, Yuqi, et al.
Published: (2025)
GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction
by: Zuo, Sicheng, et al.
Published: (2024)
by: Zuo, Sicheng, et al.
Published: (2024)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024)
by: Wu, Yuqi, et al.
Published: (2024)
SpectralAR: Spectral Autoregressive Visual Generation
by: Huang, Yuanhui, et al.
Published: (2025)
by: Huang, Yuanhui, et al.
Published: (2025)
Hardness-Aware Scene Synthesis for Semi-Supervised 3D Object Detection
by: Zeng, Shuai, et al.
Published: (2024)
by: Zeng, Shuai, et al.
Published: (2024)
Terra: Explorable Native 3D World Model with Point Latents
by: Huang, Yuanhui, et al.
Published: (2025)
by: Huang, Yuanhui, et al.
Published: (2025)
Doe-1: Closed-Loop Autonomous Driving with Large World Model
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
Driv3R: Learning Dense 4D Reconstruction for Autonomous Driving
by: Fei, Xin, et al.
Published: (2024)
by: Fei, Xin, et al.
Published: (2024)
GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
Stag-1: Towards Realistic 4D Driving Simulation with Video Generation Model
by: Wang, Lening, et al.
Published: (2024)
by: Wang, Lening, et al.
Published: (2024)
PixelGaussian: Generalizable 3D Gaussian Reconstruction from Arbitrary Views
by: Fei, Xin, et al.
Published: (2024)
by: Fei, Xin, et al.
Published: (2024)
Astra: General Interactive World Model with Autoregressive Denoising
by: Zhu, Yixuan, et al.
Published: (2025)
by: Zhu, Yixuan, et al.
Published: (2025)
Owl-1: Omni World Model for Consistent Long Video Generation
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
by: Zhuo, Dong, et al.
Published: (2026)
by: Zhuo, Dong, et al.
Published: (2026)
GPD-1: Generative Pre-training for Driving
by: Xie, Zixun, et al.
Published: (2024)
by: Xie, Zixun, et al.
Published: (2024)
DVGT: Driving Visual Geometry Transformer
by: Zuo, Sicheng, et al.
Published: (2025)
by: Zuo, Sicheng, et al.
Published: (2025)
GaussianFormer: Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
Path Choice Matters for Clear Attribution in Path Methods
by: Zhang, Borui, et al.
Published: (2024)
by: Zhang, Borui, et al.
Published: (2024)
Preventing Local Pitfalls in Vector Quantization via Optimal Transport
by: Zhang, Borui, et al.
Published: (2024)
by: Zhang, Borui, et al.
Published: (2024)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
by: Zheng, Naishan, et al.
Published: (2025)
by: Zheng, Naishan, et al.
Published: (2025)
SFTok: Bridging the Performance Gap in Discrete Tokenizers
by: Rao, Qihang, et al.
Published: (2025)
by: Rao, Qihang, et al.
Published: (2025)
Quantize-then-Rectify: Efficient VQ-VAE Training
by: Zhang, Borui, et al.
Published: (2025)
by: Zhang, Borui, et al.
Published: (2025)
Fast Shapley Value Estimation: A Unified Approach
by: Zhang, Borui, et al.
Published: (2023)
by: Zhang, Borui, et al.
Published: (2023)
VG3T: Visual Geometry Grounded Gaussian Transformer
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
GaussianToken: An Effective Image Tokenizer with 2D Gaussian Splatting
by: Dong, Jiajun, et al.
Published: (2025)
by: Dong, Jiajun, et al.
Published: (2025)
Vega: Learning to Drive with Natural Language Instructions
by: Zuo, Sicheng, et al.
Published: (2026)
by: Zuo, Sicheng, et al.
Published: (2026)
$\bf{D^3}$QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
by: Zhang, Yanran, et al.
Published: (2025)
by: Zhang, Yanran, et al.
Published: (2025)
OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
by: Wang, Lening, et al.
Published: (2024)
by: Wang, Lening, et al.
Published: (2024)
Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
by: Yang, Xiaoyu, et al.
Published: (2025)
by: Yang, Xiaoyu, et al.
Published: (2025)
SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation
by: Li, Xuewei, et al.
Published: (2023)
by: Li, Xuewei, et al.
Published: (2023)
VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution
by: Zhu, Yixuan, et al.
Published: (2026)
by: Zhu, Yixuan, et al.
Published: (2026)
Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers
by: Zheng, Shuhong, et al.
Published: (2026)
by: Zheng, Shuhong, et al.
Published: (2026)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
by: Tang, Haotian, et al.
Published: (2024)
by: Tang, Haotian, et al.
Published: (2024)
MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE
by: Zhu, Ruijie, et al.
Published: (2026)
by: Zhu, Ruijie, et al.
Published: (2026)
Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation
by: Gong, Boyang, et al.
Published: (2026)
by: Gong, Boyang, et al.
Published: (2026)
TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding
by: Guo, Wenxuan, et al.
Published: (2025)
by: Guo, Wenxuan, et al.
Published: (2025)
Learning Visual Abstract Reasoning through Dual-Stream Networks
by: Zhao, Kai, et al.
Published: (2024)
by: Zhao, Kai, et al.
Published: (2024)
Joint 3D Geometry Reconstruction and Motion Generation for 4D Synthesis from a Single Image
by: Zhang, Yanran, et al.
Published: (2025)
by: Zhang, Yanran, et al.
Published: (2025)
FabGPT: An Efficient Large Multimodal Model for Complex Wafer Defect Knowledge Queries
by: Jiang, Yuqi, et al.
Published: (2024)
by: Jiang, Yuqi, et al.
Published: (2024)
Latent Diffusion Model without Variational Autoencoder
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
Similar Items
-
Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
by: Wu, Yuqi, et al.
Published: (2025) -
GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction
by: Zuo, Sicheng, et al.
Published: (2024) -
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024) -
SpectralAR: Spectral Autoregressive Visual Generation
by: Huang, Yuanhui, et al.
Published: (2025) -
Hardness-Aware Scene Synthesis for Semi-Supervised 3D Object Detection
by: Zeng, Shuai, et al.
Published: (2024)