GPT-4 for Occlusion Order Recovery

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Saleh, Kaziwa, Rostam, Zhyar Rzgar K, Szénási, Sándor, Vámossy, Zoltán
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908561261461504
author Saleh, Kaziwa
Rostam, Zhyar Rzgar K
Szénási, Sándor
Vámossy, Zoltán
author_facet Saleh, Kaziwa
Rostam, Zhyar Rzgar K
Szénási, Sándor
Vámossy, Zoltán
contents Occlusion remains a significant challenge for current vision models to robustly interpret complex and dense real-world images and scenes. To address this limitation and to enable accurate prediction of the occlusion order relationship between objects, we propose leveraging the advanced capability of a pre-trained GPT-4 model to deduce the order. By providing a specifically designed prompt along with the input image, GPT-4 can analyze the image and generate order predictions. The response can then be parsed to construct an occlusion matrix which can be utilized in assisting with other occlusion handling tasks and image understanding. We report the results of evaluating the model on COCOA and InstaOrder datasets. The results show that by using semantic context, visual patterns, and commonsense knowledge, the model can produce more accurate order predictions. Unlike baseline methods, the model can reason about occlusion relationships in a zero-shot fashion, which requires no annotated training data and can easily be integrated into occlusion handling frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22383
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GPT-4 for Occlusion Order Recovery
Saleh, Kaziwa
Rostam, Zhyar Rzgar K
Szénási, Sándor
Vámossy, Zoltán
Computer Vision and Pattern Recognition
I.4.5
Occlusion remains a significant challenge for current vision models to robustly interpret complex and dense real-world images and scenes. To address this limitation and to enable accurate prediction of the occlusion order relationship between objects, we propose leveraging the advanced capability of a pre-trained GPT-4 model to deduce the order. By providing a specifically designed prompt along with the input image, GPT-4 can analyze the image and generate order predictions. The response can then be parsed to construct an occlusion matrix which can be utilized in assisting with other occlusion handling tasks and image understanding. We report the results of evaluating the model on COCOA and InstaOrder datasets. The results show that by using semantic context, visual patterns, and commonsense knowledge, the model can produce more accurate order predictions. Unlike baseline methods, the model can reason about occlusion relationships in a zero-shot fashion, which requires no annotated training data and can easily be integrated into occlusion handling frameworks.
title GPT-4 for Occlusion Order Recovery
topic Computer Vision and Pattern Recognition
I.4.5
url https://arxiv.org/abs/2509.22383