C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Rongchang, Feng, Zhenhua, Xu, Tianyang, Li, Linze, Wu, Xiao-Jun, Awais, Muhammad, Atito, Sara, Kittler, Josef
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913436456189952
author Li, Rongchang
Feng, Zhenhua
Xu, Tianyang
Li, Linze
Wu, Xiao-Jun
Awais, Muhammad
Atito, Sara
Kittler, Josef
author_facet Li, Rongchang
Feng, Zhenhua
Xu, Tianyang
Li, Linze
Wu, Xiao-Jun
Awais, Muhammad
Atito, Sara
Kittler, Josef
contents Compositional actions consist of dynamic (verbs) and static (objects) concepts. Humans can easily recognize unseen compositions using the learned concepts. For machines, solving such a problem requires a model to recognize unseen actions composed of previously observed verbs and objects, thus requiring so-called compositional generalization ability. To facilitate this research, we propose a novel Zero-Shot Compositional Action Recognition (ZS-CAR) task. For evaluating the task, we construct a new benchmark, Something-composition (Sth-com), based on the widely used Something-Something V2 dataset. We also propose a novel Component-to-Composition (C2C) learning method to solve the new ZS-CAR task. C2C includes an independent component learning module and a composition inference module. Last, we devise an enhanced training strategy to address the challenges of component variations between seen and unseen compositions and to handle the subtle balance between learning seen and unseen actions. The experimental results demonstrate that the proposed framework significantly surpasses the existing compositional generalization methods and sets a new state-of-the-art. The new Sth-com benchmark and code are available at https://github.com/RongchangLi/ZSCAR_C2C.
format Preprint
id arxiv_https___arxiv_org_abs_2407_06113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
Li, Rongchang
Feng, Zhenhua
Xu, Tianyang
Li, Linze
Wu, Xiao-Jun
Awais, Muhammad
Atito, Sara
Kittler, Josef
Computer Vision and Pattern Recognition
Compositional actions consist of dynamic (verbs) and static (objects) concepts. Humans can easily recognize unseen compositions using the learned concepts. For machines, solving such a problem requires a model to recognize unseen actions composed of previously observed verbs and objects, thus requiring so-called compositional generalization ability. To facilitate this research, we propose a novel Zero-Shot Compositional Action Recognition (ZS-CAR) task. For evaluating the task, we construct a new benchmark, Something-composition (Sth-com), based on the widely used Something-Something V2 dataset. We also propose a novel Component-to-Composition (C2C) learning method to solve the new ZS-CAR task. C2C includes an independent component learning module and a composition inference module. Last, we devise an enhanced training strategy to address the challenges of component variations between seen and unseen compositions and to handle the subtle balance between learning seen and unseen actions. The experimental results demonstrate that the proposed framework significantly surpasses the existing compositional generalization methods and sets a new state-of-the-art. The new Sth-com benchmark and code are available at https://github.com/RongchangLi/ZSCAR_C2C.
title C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.06113