Global-Aware Monocular Semantic Scene Completion with State Space Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shijie, Cheng, Zhongyao, Li, Rong, Li, Shuai, Gall, Juergen, Xu, Xun, Yang, Xulei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Future-Aware Interaction Network For Motion Forecasting
by: Li, Shijie, et al.
Published: (2025)
by: Li, Shijie, et al.
Published: (2025)
A Timely Survey on Vision Transformer for Deepfake Detection
by: Wang, Zhikan, et al.
Published: (2024)
by: Wang, Zhikan, et al.
Published: (2024)
Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances
by: Wang, Qirui, et al.
Published: (2026)
by: Wang, Qirui, et al.
Published: (2026)
Multi-View Industrial Anomaly Detection with Epipolar Constrained Cross-View Fusion
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Zero-Shot 3D Visual Grounding from Vision-Language Models
by: Li, Rong, et al.
Published: (2025)
by: Li, Rong, et al.
Published: (2025)
DiffPCN: Latent Diffusion Model Based on Multi-view Depth Images for Point Cloud Completion
by: Li, Zijun, et al.
Published: (2025)
by: Li, Zijun, et al.
Published: (2025)
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
by: Wang, Qirui, et al.
Published: (2025)
by: Wang, Qirui, et al.
Published: (2025)
On-the-fly Point Feature Representation for Point Clouds Analysis
by: Wang, Jiangyi, et al.
Published: (2024)
by: Wang, Jiangyi, et al.
Published: (2024)
TFNet: Exploiting Temporal Cues for Fast and Accurate LiDAR Semantic Segmentation
by: Li, Rong, et al.
Published: (2023)
by: Li, Rong, et al.
Published: (2023)
Learning a Neural Association Network for Self-supervised Multi-Object Tracking
by: Li, Shuai, et al.
Published: (2024)
by: Li, Shuai, et al.
Published: (2024)
RiverMamba: A State Space Model for Global River Discharge and Flood Forecasting
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2025)
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2025)
Perception-Aware Multimodal Spatial Reasoning from Monocular Images
by: Cheng, Yanchun, et al.
Published: (2026)
by: Cheng, Yanchun, et al.
Published: (2026)
Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion
by: Liang, Li, et al.
Published: (2025)
by: Liang, Li, et al.
Published: (2025)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
by: Gong, Zhantao, et al.
Published: (2025)
by: Gong, Zhantao, et al.
Published: (2025)
Improving Adversarial Robustness for 3D Point Cloud Recognition at Test-Time through Purified Self-Training
by: Lin, Jinpeng, et al.
Published: (2024)
by: Lin, Jinpeng, et al.
Published: (2024)
Monocular Semantic Scene Completion via Masked Recurrent Networks
by: Wang, Xuzhi, et al.
Published: (2025)
by: Wang, Xuzhi, et al.
Published: (2025)
Exploring Human-in-the-Loop Test-Time Adaptation by Synergizing Active Learning and Model Selection
by: Li, Yushu, et al.
Published: (2024)
by: Li, Yushu, et al.
Published: (2024)
GroupMamba: Efficient Group-Based Visual State Space Model
by: Shaker, Abdelrahman, et al.
Published: (2024)
by: Shaker, Abdelrahman, et al.
Published: (2024)
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
by: Cheng, Qinchuan, et al.
Published: (2026)
by: Cheng, Qinchuan, et al.
Published: (2026)
SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction
by: Pallotta, Enrico, et al.
Published: (2025)
by: Pallotta, Enrico, et al.
Published: (2025)
AdaSFormer: Adaptive Serialized Transformers for Monocular Semantic Scene Completion from Indoor Environments
by: Wang, Xuzhi, et al.
Published: (2026)
by: Wang, Xuzhi, et al.
Published: (2026)
Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score Propagation
by: Chen, Tiankai, et al.
Published: (2025)
by: Chen, Tiankai, et al.
Published: (2025)
Not All Voxels Are Equal: Hardness-Aware Semantic Scene Completion with Self-Distillation
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Fake It To Make It: Virtual Multiviews to Enhance Monocular Indoor Semantic Scene Completion
by: Selvakumar, Anith, et al.
Published: (2025)
by: Selvakumar, Anith, et al.
Published: (2025)
One Step Closer: Creating the Future to Boost Monocular Semantic Scene Completion
by: Lu, Haoang, et al.
Published: (2025)
by: Lu, Haoang, et al.
Published: (2025)
Context and Geometry Aware Voxel Transformer for Semantic Scene Completion
by: Yu, Zhu, et al.
Published: (2024)
by: Yu, Zhu, et al.
Published: (2024)
PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving
by: Pan, Yining, et al.
Published: (2026)
by: Pan, Yining, et al.
Published: (2026)
FlowSSC: Universal Generative Monocular Semantic Scene Completion via One-Step Latent Diffusion
by: Xi, Zichen, et al.
Published: (2026)
by: Xi, Zichen, et al.
Published: (2026)
3DMambaComplete: Exploring Structured State Space Model for Point Cloud Completion
by: Li, Yixuan, et al.
Published: (2024)
by: Li, Yixuan, et al.
Published: (2024)
Label-efficient Semantic Scene Completion with Scribble Annotations
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Utilizing the Mean Teacher with Supcontrast Loss for Wafer Pattern Recognition
by: Wei, Qiyu, et al.
Published: (2024)
by: Wei, Qiyu, et al.
Published: (2024)
RWKV-PCSSC: Exploring RWKV Model for Point Cloud Semantic Scene Completion
by: He, Wenzhe, et al.
Published: (2025)
by: He, Wenzhe, et al.
Published: (2025)
DepthSSC: Monocular 3D Semantic Scene Completion via Depth-Spatial Alignment and Voxel Adaptation
by: Yao, Jiawei, et al.
Published: (2023)
by: Yao, Jiawei, et al.
Published: (2023)
ET-Former: Efficient Triplane Deformable Attention for 3D Semantic Scene Completion From Monocular Camera
by: Liang, Jing, et al.
Published: (2024)
by: Liang, Jing, et al.
Published: (2024)
Robust Distribution Alignment for Industrial Anomaly Detection under Distribution Shift
by: Liao, Jingyi, et al.
Published: (2025)
by: Liao, Jingyi, et al.
Published: (2025)
VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible Regions
by: Lu, Haoang, et al.
Published: (2025)
by: Lu, Haoang, et al.
Published: (2025)
AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization
by: Liao, Jingyi, et al.
Published: (2025)
by: Liao, Jingyi, et al.
Published: (2025)
Learning Temporal 3D Semantic Scene Completion via Optical Flow Guidance
by: Wang, Meng, et al.
Published: (2025)
by: Wang, Meng, et al.
Published: (2025)
Vision-based 3D Semantic Scene Completion via Capture Dynamic Representations
by: Wang, Meng, et al.
Published: (2025)
by: Wang, Meng, et al.
Published: (2025)
Similar Items
-
Future-Aware Interaction Network For Motion Forecasting
by: Li, Shijie, et al.
Published: (2025) -
A Timely Survey on Vision Transformer for Deepfake Detection
by: Wang, Zhikan, et al.
Published: (2024) -
Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances
by: Wang, Qirui, et al.
Published: (2026) -
Multi-View Industrial Anomaly Detection with Epipolar Constrained Cross-View Fusion
by: Liu, Yifan, et al.
Published: (2025) -
Zero-Shot 3D Visual Grounding from Vision-Language Models
by: Li, Rong, et al.
Published: (2025)