Prototypical Transformer as Unified Motion Learners

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Han, Cheng, Lu, Yawen, Sun, Guohao, Liang, James C., Cao, Zhiwen, Wang, Qifan, Guan, Qiang, Dianat, Sohail A., Rao, Raghuveer M., Geng, Tong, Tao, Zhiqiang, Liu, Dongfang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914821022154752
author Han, Cheng
Lu, Yawen
Sun, Guohao
Liang, James C.
Cao, Zhiwen
Wang, Qifan
Guan, Qiang
Dianat, Sohail A.
Rao, Raghuveer M.
Geng, Tong
Tao, Zhiqiang
Liu, Dongfang
author_facet Han, Cheng
Lu, Yawen
Sun, Guohao
Liang, James C.
Cao, Zhiwen
Wang, Qifan
Guan, Qiang
Dianat, Sohail A.
Rao, Raghuveer M.
Geng, Tong
Tao, Zhiqiang
Liu, Dongfang
contents In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoFormer seamlessly integrates prototype learning with Transformer by thoughtfully considering motion dynamics, introducing two innovative designs. First, Cross-Attention Prototyping discovers prototypes based on signature motion patterns, providing transparency in understanding motion scenes. Second, Latent Synchronization guides feature representation learning via prototypes, effectively mitigating the problem of motion uncertainty. Empirical results demonstrate that our approach achieves competitive performance on popular motion tasks such as optical flow and scene depth. Furthermore, it exhibits generality across various downstream tasks, including object tracking and video stabilization.
format Preprint
id arxiv_https___arxiv_org_abs_2406_01559
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prototypical Transformer as Unified Motion Learners
Han, Cheng
Lu, Yawen
Sun, Guohao
Liang, James C.
Cao, Zhiwen
Wang, Qifan
Guan, Qiang
Dianat, Sohail A.
Rao, Raghuveer M.
Geng, Tong
Tao, Zhiqiang
Liu, Dongfang
Computer Vision and Pattern Recognition
In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoFormer seamlessly integrates prototype learning with Transformer by thoughtfully considering motion dynamics, introducing two innovative designs. First, Cross-Attention Prototyping discovers prototypes based on signature motion patterns, providing transparency in understanding motion scenes. Second, Latent Synchronization guides feature representation learning via prototypes, effectively mitigating the problem of motion uncertainty. Empirical results demonstrate that our approach achieves competitive performance on popular motion tasks such as optical flow and scene depth. Furthermore, it exhibits generality across various downstream tasks, including object tracking and video stabilization.
title Prototypical Transformer as Unified Motion Learners
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.01559