CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Jinzhou, Zhou, Jie, Xu, Wenhao, Xu, Rongtao, Wang, Changwei, Chen, Shunpeng, Fu, Kexue, Shao, Yihua, Guo, Li, Xu, Shibiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909844701708288
author Lin, Jinzhou
Zhou, Jie
Xu, Wenhao
Xu, Rongtao
Wang, Changwei
Chen, Shunpeng
Fu, Kexue
Shao, Yihua
Guo, Li
Xu, Shibiao
author_facet Lin, Jinzhou
Zhou, Jie
Xu, Wenhao
Xu, Rongtao
Wang, Changwei
Chen, Shunpeng
Fu, Kexue
Shao, Yihua
Guo, Li
Xu, Shibiao
contents Semantic Scene Completion (SSC) aims to infer complete 3D geometry and semantics from monocular images, serving as a crucial capability for camera-based perception in autonomous driving. However, existing SSC methods relying on temporal stacking or depth projection often lack explicit motion reasoning and struggle with occlusions and noisy depth supervision. We propose CurriFlow, a novel semantic occupancy prediction framework that integrates optical flow-based temporal alignment with curriculum-guided depth fusion. CurriFlow employs a multi-level fusion strategy to align segmentation, visual, and depth features across frames using pre-trained optical flow, thereby improving temporal consistency and dynamic object understanding. To enhance geometric robustness, a curriculum learning mechanism progressively transitions from sparse yet accurate LiDAR depth to dense but noisy stereo depth during training, ensuring stable optimization and seamless adaptation to real-world deployment. Furthermore, semantic priors from the Segment Anything Model (SAM) provide category-agnostic supervision, strengthening voxel-level semantic learning and spatial consistency. Experiments on the SemanticKITTI benchmark demonstrate that CurriFlow achieves state-of-the-art performance with a mean IoU of 16.9, validating the effectiveness of our motion-guided and curriculum-aware design for camera-based 3D semantic scene completion.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12362
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
Lin, Jinzhou
Zhou, Jie
Xu, Wenhao
Xu, Rongtao
Wang, Changwei
Chen, Shunpeng
Fu, Kexue
Shao, Yihua
Guo, Li
Xu, Shibiao
Computer Vision and Pattern Recognition
Semantic Scene Completion (SSC) aims to infer complete 3D geometry and semantics from monocular images, serving as a crucial capability for camera-based perception in autonomous driving. However, existing SSC methods relying on temporal stacking or depth projection often lack explicit motion reasoning and struggle with occlusions and noisy depth supervision. We propose CurriFlow, a novel semantic occupancy prediction framework that integrates optical flow-based temporal alignment with curriculum-guided depth fusion. CurriFlow employs a multi-level fusion strategy to align segmentation, visual, and depth features across frames using pre-trained optical flow, thereby improving temporal consistency and dynamic object understanding. To enhance geometric robustness, a curriculum learning mechanism progressively transitions from sparse yet accurate LiDAR depth to dense but noisy stereo depth during training, ensuring stable optimization and seamless adaptation to real-world deployment. Furthermore, semantic priors from the Segment Anything Model (SAM) provide category-agnostic supervision, strengthening voxel-level semantic learning and spatial consistency. Experiments on the SemanticKITTI benchmark demonstrate that CurriFlow achieves state-of-the-art performance with a mean IoU of 16.9, validating the effectiveness of our motion-guided and curriculum-aware design for camera-based 3D semantic scene completion.
title CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.12362