SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Dang, Chenxu, Wang, Jie, Li, Guang, Hou, Zhiwen, You, Zihan, Ye, Hangjun, Ma, Jie, Chen, Long, Wang, Yan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries
by: Dang, Chenxu, et al.
Published: (2025)
by: Dang, Chenxu, et al.
Published: (2025)
DriveFine: Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving
by: Dang, Chenxu, et al.
Published: (2026)
by: Dang, Chenxu, et al.
Published: (2026)
SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction
by: Tang, Pin, et al.
Published: (2024)
by: Tang, Pin, et al.
Published: (2024)
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
by: You, Zihan, et al.
Published: (2026)
by: You, Zihan, et al.
Published: (2026)
VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
by: Wang, Jie, et al.
Published: (2026)
by: Wang, Jie, et al.
Published: (2026)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024)
by: Wu, Yuqi, et al.
Published: (2024)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
OPUS: Occupancy Prediction Using a Sparse Set
by: Wang, Jiabao, et al.
Published: (2024)
by: Wang, Jiabao, et al.
Published: (2024)
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
by: Liu, Ruixun, et al.
Published: (2025)
by: Liu, Ruixun, et al.
Published: (2025)
OctreeOcc: Efficient and Multi-Granularity Occupancy Prediction Using Octree Queries
by: Lu, Yuhang, et al.
Published: (2023)
by: Lu, Yuhang, et al.
Published: (2023)
UniOcc: A Unified Benchmark for Occupancy Forecasting and Prediction in Autonomous Driving
by: Wang, Yuping, et al.
Published: (2025)
by: Wang, Yuping, et al.
Published: (2025)
QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy
by: Lilja, Adam, et al.
Published: (2025)
by: Lilja, Adam, et al.
Published: (2025)
SPOT-Occ: Sparse Prototype-guided Transformer for Camera-based 3D Occupancy Prediction
by: Chen, Suzeyu, et al.
Published: (2026)
by: Chen, Suzeyu, et al.
Published: (2026)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
GaussianFlowOcc: Sparse and Weakly Supervised Occupancy Estimation using Gaussian Splatting and Temporal Flow
by: Boeder, Simon, et al.
Published: (2025)
by: Boeder, Simon, et al.
Published: (2025)
SUG-Occ: Explicit Semantics and Uncertainty Guided Sparse Learning for Efficient 3D Occupancy Prediction
by: Wu, Hanlin, et al.
Published: (2026)
by: Wu, Hanlin, et al.
Published: (2026)
ForecastOcc: Vision-based Semantic Occupancy Forecasting
by: Mohan, Riya, et al.
Published: (2026)
by: Mohan, Riya, et al.
Published: (2026)
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
OccSTeP: Benchmarking 4D Occupancy Spatio-Temporal Persistence
by: Zheng, Yu, et al.
Published: (2025)
by: Zheng, Yu, et al.
Published: (2025)
MergeOcc: Bridge the Domain Gap between Different LiDARs for Robust Occupancy Prediction
by: Xu, Zikun, et al.
Published: (2024)
by: Xu, Zikun, et al.
Published: (2024)
SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
by: Du, Jiayuan, et al.
Published: (2025)
by: Du, Jiayuan, et al.
Published: (2025)
AdaOcc: Adaptive-Resolution Occupancy Prediction
by: Chen, Chao, et al.
Published: (2024)
by: Chen, Chao, et al.
Published: (2024)
GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting
by: Ye, Baijun, et al.
Published: (2025)
by: Ye, Baijun, et al.
Published: (2025)
OccMamba: Semantic Occupancy Prediction with State Space Models
by: Li, Heng, et al.
Published: (2024)
by: Li, Heng, et al.
Published: (2024)
MonoOcc: Digging into Monocular Semantic Occupancy Prediction
by: Zheng, Yupeng, et al.
Published: (2024)
by: Zheng, Yupeng, et al.
Published: (2024)
OpenOcc: Open Vocabulary 3D Scene Reconstruction via Occupancy Representation
by: Jiang, Haochen, et al.
Published: (2024)
by: Jiang, Haochen, et al.
Published: (2024)
STCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Prediction
by: Liao, Zhimin, et al.
Published: (2025)
by: Liao, Zhimin, et al.
Published: (2025)
Fully Sparse 3D Occupancy Prediction
by: Liu, Haisong, et al.
Published: (2023)
by: Liu, Haisong, et al.
Published: (2023)
UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
by: Tan, Kaiyuan, et al.
Published: (2026)
by: Tan, Kaiyuan, et al.
Published: (2026)
SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction
by: Duan, Zaipeng, et al.
Published: (2025)
by: Duan, Zaipeng, et al.
Published: (2025)
OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL
by: Jie, Haoxiang, et al.
Published: (2026)
by: Jie, Haoxiang, et al.
Published: (2026)
RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
by: Pan, Mingjie, et al.
Published: (2023)
by: Pan, Mingjie, et al.
Published: (2023)
OccRWKV: Rethinking Efficient 3D Semantic Occupancy Prediction with Linear Complexity
by: Wang, Junming, et al.
Published: (2024)
by: Wang, Junming, et al.
Published: (2024)
Towards Sparse Video Understanding and Reasoning
by: Xu, Chenwei, et al.
Published: (2026)
by: Xu, Chenwei, et al.
Published: (2026)
AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting
by: Zhou, Xiaoyu, et al.
Published: (2025)
by: Zhou, Xiaoyu, et al.
Published: (2025)
CVT-Occ: Cost Volume Temporal Fusion for 3D Occupancy Prediction
by: Ye, Zhangchen, et al.
Published: (2024)
by: Ye, Zhangchen, et al.
Published: (2024)
A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse Signals
by: Tang, Jiangnan, et al.
Published: (2024)
by: Tang, Jiangnan, et al.
Published: (2024)
PanopticQuery: Unified Query-Time Reasoning for 4D Scenes
by: Tang, Ruilin, et al.
Published: (2026)
by: Tang, Ruilin, et al.
Published: (2026)
ReliOcc: Towards Reliable Semantic Occupancy Prediction via Uncertainty Learning
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
by: Luo, Yuechen, et al.
Published: (2026)
by: Luo, Yuechen, et al.
Published: (2026)
Similar Items
-
SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries
by: Dang, Chenxu, et al.
Published: (2025) -
DriveFine: Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving
by: Dang, Chenxu, et al.
Published: (2026) -
SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction
by: Tang, Pin, et al.
Published: (2024) -
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
by: You, Zihan, et al.
Published: (2026) -
VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
by: Wang, Jie, et al.
Published: (2026)