1DFormer: a Transformer Architecture Learning 1D Landmark Representations for Facial Landmark Tracking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Shi, Huan, Shijie, Wang, Shangfei, Hu, Jinshui, Guo, Tao, Yin, Bing, Yin, Baocai, Liu, Cong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909089675608064
author Yin, Shi
Huan, Shijie
Wang, Shangfei
Hu, Jinshui
Guo, Tao
Yin, Bing
Yin, Baocai
Liu, Cong
author_facet Yin, Shi
Huan, Shijie
Wang, Shangfei
Hu, Jinshui
Guo, Tao
Yin, Bing
Yin, Baocai
Liu, Cong
contents Recently, heatmap regression methods based on 1D landmark representations have shown prominent performance on locating facial landmarks. However, previous methods ignored to make deep explorations on the good potentials of 1D landmark representations for sequential and structural modeling of multiple landmarks to track facial landmarks. To address this limitation, we propose a Transformer architecture, namely 1DFormer, which learns informative 1D landmark representations by capturing the dynamic and the geometric patterns of landmarks via token communications in both temporal and spatial dimensions for facial landmark tracking. For temporal modeling, we propose a recurrent token mixing mechanism, an axis-landmark-positional embedding mechanism, as well as a confidence-enhanced multi-head attention mechanism to adaptively and robustly embed long-term landmark dynamics into their 1D representations; for structure modeling, we design intra-group and inter-group structure modeling mechanisms to encode the component-level as well as global-level facial structure patterns as a refinement for the 1D representations of landmarks through token communications in the spatial dimension via 1D convolutional layers. Experimental results on the 300VW and the TF databases show that 1DFormer successfully models the long-range sequential patterns as well as the inherent facial structures to learn informative 1D representations of landmark sequences, and achieves state-of-the-art performance on facial landmark tracking.
format Preprint
id arxiv_https___arxiv_org_abs_2311_00241
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle 1DFormer: a Transformer Architecture Learning 1D Landmark Representations for Facial Landmark Tracking
Yin, Shi
Huan, Shijie
Wang, Shangfei
Hu, Jinshui
Guo, Tao
Yin, Bing
Yin, Baocai
Liu, Cong
Computer Vision and Pattern Recognition
Recently, heatmap regression methods based on 1D landmark representations have shown prominent performance on locating facial landmarks. However, previous methods ignored to make deep explorations on the good potentials of 1D landmark representations for sequential and structural modeling of multiple landmarks to track facial landmarks. To address this limitation, we propose a Transformer architecture, namely 1DFormer, which learns informative 1D landmark representations by capturing the dynamic and the geometric patterns of landmarks via token communications in both temporal and spatial dimensions for facial landmark tracking. For temporal modeling, we propose a recurrent token mixing mechanism, an axis-landmark-positional embedding mechanism, as well as a confidence-enhanced multi-head attention mechanism to adaptively and robustly embed long-term landmark dynamics into their 1D representations; for structure modeling, we design intra-group and inter-group structure modeling mechanisms to encode the component-level as well as global-level facial structure patterns as a refinement for the 1D representations of landmarks through token communications in the spatial dimension via 1D convolutional layers. Experimental results on the 300VW and the TF databases show that 1DFormer successfully models the long-range sequential patterns as well as the inherent facial structures to learn informative 1D representations of landmark sequences, and achieves state-of-the-art performance on facial landmark tracking.
title 1DFormer: a Transformer Architecture Learning 1D Landmark Representations for Facial Landmark Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.00241