Guardado en:
Detalles Bibliográficos
Autores principales: Yeo, Qi Xun, Li, Yanyan, Lee, Gim Hee
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2508.06546
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915436761710592
author Yeo, Qi Xun
Li, Yanyan
Lee, Gim Hee
author_facet Yeo, Qi Xun
Li, Yanyan
Lee, Gim Hee
contents Modern 3D semantic scene graph estimation methods utilize ground truth 3D annotations to accurately predict target objects, predicates, and relationships. In the absence of given 3D ground truth representations, we explore leveraging only multi-view RGB images to tackle this task. To attain robust features for accurate scene graph estimation, we must overcome the noisy reconstructed pseudo point-based geometry from predicted depth maps and reduce the amount of background noise present in multi-view image features. The key is to enrich node and edge features with accurate semantic and spatial information and through neighboring relations. We obtain semantic masks to guide feature aggregation to filter background features and design a novel method to incorporate neighboring node information to aid robustness of our scene graph estimates. Furthermore, we leverage on explicit statistical priors calculated from the training summary statistics to refine node and edge predictions based on their one-hop neighborhood. Our experiments show that our method outperforms current methods purely using multi-view images as the initial input. Our project page is available at https://qixun1.github.io/projects/SCRSSG.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06546
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View Images
Yeo, Qi Xun
Li, Yanyan
Lee, Gim Hee
Computer Vision and Pattern Recognition
Image and Video Processing
Modern 3D semantic scene graph estimation methods utilize ground truth 3D annotations to accurately predict target objects, predicates, and relationships. In the absence of given 3D ground truth representations, we explore leveraging only multi-view RGB images to tackle this task. To attain robust features for accurate scene graph estimation, we must overcome the noisy reconstructed pseudo point-based geometry from predicted depth maps and reduce the amount of background noise present in multi-view image features. The key is to enrich node and edge features with accurate semantic and spatial information and through neighboring relations. We obtain semantic masks to guide feature aggregation to filter background features and design a novel method to incorporate neighboring node information to aid robustness of our scene graph estimates. Furthermore, we leverage on explicit statistical priors calculated from the training summary statistics to refine node and edge predictions based on their one-hop neighborhood. Our experiments show that our method outperforms current methods purely using multi-view images as the initial input. Our project page is available at https://qixun1.github.io/projects/SCRSSG.
title Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View Images
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2508.06546