Saved in:
Bibliographic Details
Main Authors: Liang, Guoqiang, Hu, Jiahao, Wang, Qingyue, Zhang, Shizhou
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2402.04558
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917584143646720
author Liang, Guoqiang
Hu, Jiahao
Wang, Qingyue
Zhang, Shizhou
author_facet Liang, Guoqiang
Hu, Jiahao
Wang, Qingyue
Zhang, Shizhou
contents Human de-occlusion, which aims to infer the appearance of invisible human parts from an occluded image, has great value in many human-related tasks, such as person re-id, and intention inference. To address this task, this paper proposes a dynamic mask-aware transformer (DMAT), which dynamically augments information from human regions and weakens that from occlusion. First, to enhance token representation, we design an expanded convolution head with enlarged kernels, which captures more local valid context and mitigates the influence of surrounding occlusion. To concentrate on the visible human parts, we propose a novel dynamic multi-head human-mask guided attention mechanism through integrating multiple masks, which can prevent the de-occluded regions from assimilating to the background. Besides, a region upsampling strategy is utilized to alleviate the impact of occlusion on interpolated images. During model learning, an amodal loss is developed to further emphasize the recovery effect of human regions, which also refines the model's convergence. Extensive experiments on the AHP dataset demonstrate its superior performance compared to recent state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2402_04558
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DMAT: A Dynamic Mask-Aware Transformer for Human De-occlusion
Liang, Guoqiang
Hu, Jiahao
Wang, Qingyue
Zhang, Shizhou
Computer Vision and Pattern Recognition
Human de-occlusion, which aims to infer the appearance of invisible human parts from an occluded image, has great value in many human-related tasks, such as person re-id, and intention inference. To address this task, this paper proposes a dynamic mask-aware transformer (DMAT), which dynamically augments information from human regions and weakens that from occlusion. First, to enhance token representation, we design an expanded convolution head with enlarged kernels, which captures more local valid context and mitigates the influence of surrounding occlusion. To concentrate on the visible human parts, we propose a novel dynamic multi-head human-mask guided attention mechanism through integrating multiple masks, which can prevent the de-occluded regions from assimilating to the background. Besides, a region upsampling strategy is utilized to alleviate the impact of occlusion on interpolated images. During model learning, an amodal loss is developed to further emphasize the recovery effect of human regions, which also refines the model's convergence. Extensive experiments on the AHP dataset demonstrate its superior performance compared to recent state-of-the-art methods.
title DMAT: A Dynamic Mask-Aware Transformer for Human De-occlusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.04558