Edge Prediction for Roof Wireframe Reconstruction with Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hanning, Gustav, Dillén, Ludvig, Astermark, Jonathan, Lidholm, Johanna, Larsson, Viktor
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916075043553280
author Hanning, Gustav
Dillén, Ludvig
Astermark, Jonathan
Lidholm, Johanna
Larsson, Viktor
author_facet Hanning, Gustav
Dillén, Ludvig
Astermark, Jonathan
Lidholm, Johanna
Larsson, Viktor
contents This paper presents a competitive solution to the S23DR Challenge 2026, which aims to reconstruct 3D house roof wireframe models from sparse SfM point clouds and ground-level semantic segmentations and depth maps. Our proposed method utilizes an end-to-end Transformer encoder-decoder architecture inspired by DETR. To effectively process the geometric and semantic data, the sparse SfM point cloud input is dynamically subsampled based on semantic priority and augmented with Gestalt and ADE20k class features. To further increase segmentation context, we fuse the point features with additional Gestalt feature encodings which are obtained by projecting the points into latent feature maps produced by a frozen autoencoder. Learned query embeddings are then decoded directly into 3D wireframe edges via cross-attention mechanisms. Evaluated on the "HoHo 22k" dataset, our approach significantly outperforms both handcrafted and learned baselines, achieving a Hybrid Structure Score (HSS) of 0.6476 and securing the second-highest position on the challenge's private leaderboard.
format Preprint
id arxiv_https___arxiv_org_abs_2606_02406
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Edge Prediction for Roof Wireframe Reconstruction with Transformers
Hanning, Gustav
Dillén, Ludvig
Astermark, Jonathan
Lidholm, Johanna
Larsson, Viktor
Computer Vision and Pattern Recognition
This paper presents a competitive solution to the S23DR Challenge 2026, which aims to reconstruct 3D house roof wireframe models from sparse SfM point clouds and ground-level semantic segmentations and depth maps. Our proposed method utilizes an end-to-end Transformer encoder-decoder architecture inspired by DETR. To effectively process the geometric and semantic data, the sparse SfM point cloud input is dynamically subsampled based on semantic priority and augmented with Gestalt and ADE20k class features. To further increase segmentation context, we fuse the point features with additional Gestalt feature encodings which are obtained by projecting the points into latent feature maps produced by a frozen autoencoder. Learned query embeddings are then decoded directly into 3D wireframe edges via cross-attention mechanisms. Evaluated on the "HoHo 22k" dataset, our approach significantly outperforms both handcrafted and learned baselines, achieving a Hybrid Structure Score (HSS) of 0.6476 and securing the second-highest position on the challenge's private leaderboard.
title Edge Prediction for Roof Wireframe Reconstruction with Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2606.02406