Transformer-Based Inpainting for Real-Time 3D Streaming in Sparse Multi-Camera Setups

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Van Holland, Leif, Zingsheim, Domenic, Takhsha, Mana, Dröge, Hannah, Stotko, Patrick, Plack, Markus, Klein, Reinhard
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912945797070848
author Van Holland, Leif
Zingsheim, Domenic
Takhsha, Mana
Dröge, Hannah
Stotko, Patrick
Plack, Markus
Klein, Reinhard
author_facet Van Holland, Leif
Zingsheim, Domenic
Takhsha, Mana
Dröge, Hannah
Stotko, Patrick
Plack, Markus
Klein, Reinhard
contents High-quality 3D streaming from multiple cameras is crucial for immersive experiences in many AR/VR applications. The limited number of views - often due to real-time constraints - leads to missing information and incomplete surfaces in the rendered images. Existing approaches typically rely on simple heuristics for the hole filling, which can result in inconsistencies or visual artifacts. We propose to complete the missing textures using a novel, application-targeted inpainting method independent of the underlying representation as an image-based post-processing step after the novel view rendering. The method is designed as a standalone module compatible with any calibrated multi-camera system. For this we introduce a multi-view aware, transformer-based network architecture using spatio-temporal embeddings to ensure consistency across frames while preserving fine details. Additionally, our resolution-independent design allows adaptation to different camera setups, while an adaptive patch selection strategy balances inference speed and quality, allowing real-time performance. We evaluate our approach against state-of-the-art inpainting techniques under the same real-time constraints and demonstrate that our model achieves the best trade-off between quality and speed, outperforming competitors in both image and video-based metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05507
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Transformer-Based Inpainting for Real-Time 3D Streaming in Sparse Multi-Camera Setups
Van Holland, Leif
Zingsheim, Domenic
Takhsha, Mana
Dröge, Hannah
Stotko, Patrick
Plack, Markus
Klein, Reinhard
Computer Vision and Pattern Recognition
Graphics
High-quality 3D streaming from multiple cameras is crucial for immersive experiences in many AR/VR applications. The limited number of views - often due to real-time constraints - leads to missing information and incomplete surfaces in the rendered images. Existing approaches typically rely on simple heuristics for the hole filling, which can result in inconsistencies or visual artifacts. We propose to complete the missing textures using a novel, application-targeted inpainting method independent of the underlying representation as an image-based post-processing step after the novel view rendering. The method is designed as a standalone module compatible with any calibrated multi-camera system. For this we introduce a multi-view aware, transformer-based network architecture using spatio-temporal embeddings to ensure consistency across frames while preserving fine details. Additionally, our resolution-independent design allows adaptation to different camera setups, while an adaptive patch selection strategy balances inference speed and quality, allowing real-time performance. We evaluate our approach against state-of-the-art inpainting techniques under the same real-time constraints and demonstrate that our model achieves the best trade-off between quality and speed, outperforming competitors in both image and video-based metrics.
title Transformer-Based Inpainting for Real-Time 3D Streaming in Sparse Multi-Camera Setups
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2603.05507