Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhixuan, Yoon, Hyunse, Lee, Sanghoon, Lin, Weisi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918092346490880
author Li, Zhixuan
Yoon, Hyunse
Lee, Sanghoon
Lin, Weisi
author_facet Li, Zhixuan
Yoon, Hyunse
Lee, Sanghoon
Lin, Weisi
contents Amodal segmentation aims to infer the complete shape of occluded objects, even when the occluded region's appearance is unavailable. However, current amodal segmentation methods lack the capability to interact with users through text input and struggle to understand or reason about implicit and complex purposes. While methods like LISA integrate multi-modal large language models (LLMs) with segmentation for reasoning tasks, they are limited to predicting only visible object regions and face challenges in handling complex occlusion scenarios. To address these limitations, we propose a novel task named amodal reasoning segmentation, aiming to predict the complete amodal shape of occluded objects while providing answers with elaborations based on user text input. We develop a generalizable dataset generation pipeline and introduce a new dataset focusing on daily life scenarios, encompassing diverse real-world occlusions. Furthermore, we present AURA (Amodal Understanding and Reasoning Assistant), a novel model with advanced global and spatial-level designs specifically tailored to handle complex occlusions. Extensive experiments validate AURA's effectiveness on the proposed dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10225
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA
Li, Zhixuan
Yoon, Hyunse
Lee, Sanghoon
Lin, Weisi
Computer Vision and Pattern Recognition
Amodal segmentation aims to infer the complete shape of occluded objects, even when the occluded region's appearance is unavailable. However, current amodal segmentation methods lack the capability to interact with users through text input and struggle to understand or reason about implicit and complex purposes. While methods like LISA integrate multi-modal large language models (LLMs) with segmentation for reasoning tasks, they are limited to predicting only visible object regions and face challenges in handling complex occlusion scenarios. To address these limitations, we propose a novel task named amodal reasoning segmentation, aiming to predict the complete amodal shape of occluded objects while providing answers with elaborations based on user text input. We develop a generalizable dataset generation pipeline and introduce a new dataset focusing on daily life scenarios, encompassing diverse real-world occlusions. Furthermore, we present AURA (Amodal Understanding and Reasoning Assistant), a novel model with advanced global and spatial-level designs specifically tailored to handle complex occlusions. Extensive experiments validate AURA's effectiveness on the proposed dataset.
title Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.10225