Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Calabrese, Carmela, Berti, Stefano, Pasquale, Giulia, Natale, Lorenzo
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2405.08695
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910446744764416
author Calabrese, Carmela
Berti, Stefano
Pasquale, Giulia
Natale, Lorenzo
author_facet Calabrese, Carmela
Berti, Stefano
Pasquale, Giulia
Natale, Lorenzo
contents Addressing multi-label action recognition in videos represents a significant challenge for robotic applications in dynamic environments, especially when the robot is required to cooperate with humans in tasks that involve objects. Existing methods still struggle to recognize unseen actions or require extensive training data. To overcome these problems, we propose Dual-VCLIP, a unified approach for zero-shot multi-label action recognition. Dual-VCLIP enhances VCLIP, a zero-shot action recognition method, with the DualCoOp method for multi-label image classification. The strength of our method is that at training time it only learns two prompts, and it is therefore much simpler than other methods. We validate our method on the Charades dataset that includes a majority of object-based actions, demonstrating that -- despite its simplicity -- our method performs favorably with respect to existing methods on the complete dataset, and promising performance when tested on unseen actions. Our contribution emphasizes the impact of verb-object class-splits during robots' training for new cooperative tasks, highlighting the influence on the performance and giving insights into mitigating biases.
format Preprint
id arxiv_https___arxiv_org_abs_2405_08695
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The impact of Compositionality in Zero-shot Multi-label action recognition for Object-based tasks
Calabrese, Carmela
Berti, Stefano
Pasquale, Giulia
Natale, Lorenzo
Computer Vision and Pattern Recognition
Robotics
Addressing multi-label action recognition in videos represents a significant challenge for robotic applications in dynamic environments, especially when the robot is required to cooperate with humans in tasks that involve objects. Existing methods still struggle to recognize unseen actions or require extensive training data. To overcome these problems, we propose Dual-VCLIP, a unified approach for zero-shot multi-label action recognition. Dual-VCLIP enhances VCLIP, a zero-shot action recognition method, with the DualCoOp method for multi-label image classification. The strength of our method is that at training time it only learns two prompts, and it is therefore much simpler than other methods. We validate our method on the Charades dataset that includes a majority of object-based actions, demonstrating that -- despite its simplicity -- our method performs favorably with respect to existing methods on the complete dataset, and promising performance when tested on unseen actions. Our contribution emphasizes the impact of verb-object class-splits during robots' training for new cooperative tasks, highlighting the influence on the performance and giving insights into mitigating biases.
title The impact of Compositionality in Zero-shot Multi-label action recognition for Object-based tasks
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2405.08695