Language and Geometry Grounded Sparse Voxel Representations for Holistic Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Guile, Huang, David, Liu, Bingbing, Bai, Dongfeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
by: Wu, Guile, et al.
Published: (2025)
by: Wu, Guile, et al.
Published: (2025)
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
by: Huang, David, et al.
Published: (2026)
by: Huang, David, et al.
Published: (2026)
ArmGS: Composite Gaussian Appearance Refinement for Modeling Dynamic Urban Environments
by: Wu, Guile, et al.
Published: (2025)
by: Wu, Guile, et al.
Published: (2025)
Nighttime Autonomous Driving Scene Reconstruction with Physically-Based Gaussian Splatting
by: Kim, Tae-Kyeong, et al.
Published: (2026)
by: Kim, Tae-Kyeong, et al.
Published: (2026)
HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting
by: Zhou, Hongyu, et al.
Published: (2024)
by: Zhou, Hongyu, et al.
Published: (2024)
UniVoxel: Fast Inverse Rendering by Unified Voxelization of Scene Representation
by: Wu, Shuang, et al.
Published: (2024)
by: Wu, Shuang, et al.
Published: (2024)
UniGaussian: Driving Scene Reconstruction from Multiple Camera Models via Unified Gaussian Representations
by: Ren, Yuan, et al.
Published: (2024)
by: Ren, Yuan, et al.
Published: (2024)
Learning Effective NeRFs and SDFs Representations with 3D Generative Adversarial Networks for 3D Object Generation
by: Yang, Zheyuan, et al.
Published: (2023)
by: Yang, Zheyuan, et al.
Published: (2023)
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
by: Huang, Sheng-Yu, et al.
Published: (2026)
by: Huang, Sheng-Yu, et al.
Published: (2026)
Context and Geometry Aware Voxel Transformer for Semantic Scene Completion
by: Yu, Zhu, et al.
Published: (2024)
by: Yu, Zhu, et al.
Published: (2024)
UniScale: Unified Scale-Aware 3D Reconstruction for Multi-View Understanding via Prior Injection for Robotic Perception
by: Mahdavian, Mohammad, et al.
Published: (2026)
by: Mahdavian, Mohammad, et al.
Published: (2026)
EVolSplat4D: Efficient Volume-based Gaussian Splatting for 4D Urban Scene Synthesis
by: Miao, Sheng, et al.
Published: (2026)
by: Miao, Sheng, et al.
Published: (2026)
AutoSplat: Constrained Gaussian Splatting for Autonomous Driving Scene Reconstruction
by: Khan, Mustafa, et al.
Published: (2024)
by: Khan, Mustafa, et al.
Published: (2024)
VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding
by: Wang, Ruoyu, et al.
Published: (2026)
by: Wang, Ruoyu, et al.
Published: (2026)
4D Neural Voxel Splatting: Dynamic Scene Rendering with Voxelized Guassian Splatting
by: Wu, Chun-Tin, et al.
Published: (2025)
by: Wu, Chun-Tin, et al.
Published: (2025)
Dynamic Scene Understanding through Object-Centric Voxelization and Neural Rendering
by: Zhao, Yanpeng, et al.
Published: (2024)
by: Zhao, Yanpeng, et al.
Published: (2024)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
by: Ropero, Fernando, et al.
Published: (2026)
by: Ropero, Fernando, et al.
Published: (2026)
Monocular Visual 8D Pose Estimation for Articulated Bicycles and Cyclists
by: Corral-Soto, Eduardo R., et al.
Published: (2025)
by: Corral-Soto, Eduardo R., et al.
Published: (2025)
AVS-Net: Point Sampling with Adaptive Voxel Size for 3D Scene Understanding
by: Yang, Hongcheng, et al.
Published: (2024)
by: Yang, Hongcheng, et al.
Published: (2024)
Towards Holistic Surgical Scene Understanding
by: Valderrama, Natalia, et al.
Published: (2022)
by: Valderrama, Natalia, et al.
Published: (2022)
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving
by: Liu, Yibo, et al.
Published: (2024)
by: Liu, Yibo, et al.
Published: (2024)
$\text{VG}^2$GT: Voxel-Gaussian Splatting Visual Geometry Grounded Transformer
by: Zhao, Yibin, et al.
Published: (2026)
by: Zhao, Yibin, et al.
Published: (2026)
Self-Supervised Scene Flow Estimation with Point-Voxel Fusion and Surface Representation
by: Xiang, Xuezhi, et al.
Published: (2024)
by: Xiang, Xuezhi, et al.
Published: (2024)
Efficient Depth-Guided Urban View Synthesis
by: Miao, Sheng, et al.
Published: (2024)
by: Miao, Sheng, et al.
Published: (2024)
LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
by: Wang, Fusang, et al.
Published: (2026)
by: Wang, Fusang, et al.
Published: (2026)
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
by: Wu, Xianfeng, et al.
Published: (2025)
by: Wu, Xianfeng, et al.
Published: (2025)
GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction
by: Li, Jiahe, et al.
Published: (2025)
by: Li, Jiahe, et al.
Published: (2025)
Advancing Structured Priors for Sparse-Voxel Surface Reconstruction
by: Chi, Ting-Hsun, et al.
Published: (2026)
by: Chi, Ting-Hsun, et al.
Published: (2026)
AutoScape: Geometry-Consistent Long-Horizon Scene Generation
by: Chen, Jiacheng, et al.
Published: (2025)
by: Chen, Jiacheng, et al.
Published: (2025)
SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
by: Sun, Wenchao, et al.
Published: (2024)
by: Sun, Wenchao, et al.
Published: (2024)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
by: Halacheva, Anna-Maria, et al.
Published: (2024)
by: Halacheva, Anna-Maria, et al.
Published: (2024)
VoxScene: Anchor-Conditioned Voxel Diffusion for Indoor Scene Arrangement
by: Mao, Haotian, et al.
Published: (2026)
by: Mao, Haotian, et al.
Published: (2026)
FreeFix: Boosting 3D Gaussian Splatting via Fine-Tuning-Free Diffusion Models
by: Zhou, Hongyu, et al.
Published: (2026)
by: Zhou, Hongyu, et al.
Published: (2026)
Neural Radiance Fields with Torch Units
by: Ni, Bingnan, et al.
Published: (2024)
by: Ni, Bingnan, et al.
Published: (2024)
SVRecon: Sparse Voxel Rasterization for Surface Reconstruction
by: Oh, Seunghun, et al.
Published: (2025)
by: Oh, Seunghun, et al.
Published: (2025)
PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
by: Huang, Zheng, et al.
Published: (2025)
by: Huang, Zheng, et al.
Published: (2025)
SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
by: Wang, Ruoyu, et al.
Published: (2026)
by: Wang, Ruoyu, et al.
Published: (2026)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025)
by: Karanfil, Enes, et al.
Published: (2025)
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
Semi-Supervised Crowd Counting with Contextual Modeling: Facilitating Holistic Understanding of Crowd Scenes
by: Qian, Yifei, et al.
Published: (2023)
by: Qian, Yifei, et al.
Published: (2023)
Similar Items
-
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
by: Wu, Guile, et al.
Published: (2025) -
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
by: Huang, David, et al.
Published: (2026) -
ArmGS: Composite Gaussian Appearance Refinement for Modeling Dynamic Urban Environments
by: Wu, Guile, et al.
Published: (2025) -
Nighttime Autonomous Driving Scene Reconstruction with Physically-Based Gaussian Splatting
by: Kim, Tae-Kyeong, et al.
Published: (2026) -
HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting
by: Zhou, Hongyu, et al.
Published: (2024)