Answerability Fields: Answerable Location Estimation via Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Azuma, Daichi, Miyanishi, Taiki, Kurita, Shuhei, Sakamoto, Koya, Kawanabe, Motoaki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914888270479360
author Azuma, Daichi
Miyanishi, Taiki
Kurita, Shuhei
Sakamoto, Koya
Kawanabe, Motoaki
author_facet Azuma, Daichi
Miyanishi, Taiki
Kurita, Shuhei
Sakamoto, Koya
Kawanabe, Motoaki
contents In an era characterized by advancements in artificial intelligence and robotics, enabling machines to interact with and understand their environment is a critical research endeavor. In this paper, we propose Answerability Fields, a novel approach to predicting answerability within complex indoor environments. Leveraging a 3D question answering dataset, we construct a comprehensive Answerability Fields dataset, encompassing diverse scenes and questions from ScanNet. Using a diffusion model, we successfully infer and evaluate these Answerability Fields, demonstrating the importance of objects and their locations in answering questions within a scene. Our results showcase the efficacy of Answerability Fields in guiding scene-understanding tasks, laying the foundation for their application in enhancing interactions between intelligent agents and their environments.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18497
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Answerability Fields: Answerable Location Estimation via Diffusion Models
Azuma, Daichi
Miyanishi, Taiki
Kurita, Shuhei
Sakamoto, Koya
Kawanabe, Motoaki
Computer Vision and Pattern Recognition
In an era characterized by advancements in artificial intelligence and robotics, enabling machines to interact with and understand their environment is a critical research endeavor. In this paper, we propose Answerability Fields, a novel approach to predicting answerability within complex indoor environments. Leveraging a 3D question answering dataset, we construct a comprehensive Answerability Fields dataset, encompassing diverse scenes and questions from ScanNet. Using a diffusion model, we successfully infer and evaluate these Answerability Fields, demonstrating the importance of objects and their locations in answering questions within a scene. Our results showcase the efficacy of Answerability Fields in guiding scene-understanding tasks, laying the foundation for their application in enhancing interactions between intelligent agents and their environments.
title Answerability Fields: Answerable Location Estimation via Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.18497