Supplementing Missing Visions via Dialog for Scene Graph Generations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhao, Zhenghao, Zhu, Ye, Zhu, Xiaoguang, Shang, Yuzhang, Yan, Yan
Format: Preprint
Publié: 2022
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929299300286464
author Zhao, Zhenghao
Zhu, Ye
Zhu, Xiaoguang
Shang, Yuzhang
Yan, Yan
author_facet Zhao, Zhenghao
Zhu, Ye
Zhu, Xiaoguang
Shang, Yuzhang
Yan, Yan
contents Most current AI systems rely on the premise that the input visual data are sufficient to achieve competitive performance in various computer vision tasks. However, the classic task setup rarely considers the challenging, yet common practical situations where the complete visual data may be inaccessible due to various reasons (e.g., restricted view range and occlusions). To this end, we investigate a computer vision task setting with incomplete visual input data. Specifically, we exploit the Scene Graph Generation (SGG) task with various levels of visual data missingness as input. While insufficient visual input intuitively leads to performance drop, we propose to supplement the missing visions via the natural language dialog interactions to better accomplish the task objective. We design a model-agnostic Supplementary Interactive Dialog (SI-Dial) framework that can be jointly learned with most existing models, endowing the current AI systems with the ability of question-answer interactions in natural language. We demonstrate the feasibility of such a task setting with missing visual input and the effectiveness of our proposed dialog module as the supplementary information source through extensive experiments and analysis, by achieving promising performance improvement over multiple baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2204_11143
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Supplementing Missing Visions via Dialog for Scene Graph Generations
Zhao, Zhenghao
Zhu, Ye
Zhu, Xiaoguang
Shang, Yuzhang
Yan, Yan
Computer Vision and Pattern Recognition
Most current AI systems rely on the premise that the input visual data are sufficient to achieve competitive performance in various computer vision tasks. However, the classic task setup rarely considers the challenging, yet common practical situations where the complete visual data may be inaccessible due to various reasons (e.g., restricted view range and occlusions). To this end, we investigate a computer vision task setting with incomplete visual input data. Specifically, we exploit the Scene Graph Generation (SGG) task with various levels of visual data missingness as input. While insufficient visual input intuitively leads to performance drop, we propose to supplement the missing visions via the natural language dialog interactions to better accomplish the task objective. We design a model-agnostic Supplementary Interactive Dialog (SI-Dial) framework that can be jointly learned with most existing models, endowing the current AI systems with the ability of question-answer interactions in natural language. We demonstrate the feasibility of such a task setting with missing visual input and the effectiveness of our proposed dialog module as the supplementary information source through extensive experiments and analysis, by achieving promising performance improvement over multiple baselines.
title Supplementing Missing Visions via Dialog for Scene Graph Generations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2204.11143