Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Keat, Ee Yeo, Hao, Zhang, Matyasko, Alexander, Fernando, Basura
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916372291780608
author Keat, Ee Yeo
Hao, Zhang
Matyasko, Alexander
Fernando, Basura
author_facet Keat, Ee Yeo
Hao, Zhang
Matyasko, Alexander
Fernando, Basura
contents We introduce VidTFS, a Training-free, open-vocabulary video goal and action inference framework that combines the frozen vision foundational model (VFM) and large language model (LLM) with a novel dynamic Frame Selection module. Our experiments demonstrate that the proposed frame selection module improves the performance of the framework significantly. We validate the performance of the proposed VidTFS on four widely used video datasets, including CrossTask, COIN, UCF101, and ActivityNet, covering goal inference and action recognition tasks under open-vocabulary settings without requiring any training or fine-tuning. The results show that VidTFS outperforms pretrained and instruction-tuned multimodal language models that directly stack LLM and VFM for downstream video inference tasks. Our VidTFS with its adaptability shows the future potential for generalizing to new training-free video inference tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12471
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
Keat, Ee Yeo
Hao, Zhang
Matyasko, Alexander
Fernando, Basura
Computer Vision and Pattern Recognition
We introduce VidTFS, a Training-free, open-vocabulary video goal and action inference framework that combines the frozen vision foundational model (VFM) and large language model (LLM) with a novel dynamic Frame Selection module. Our experiments demonstrate that the proposed frame selection module improves the performance of the framework significantly. We validate the performance of the proposed VidTFS on four widely used video datasets, including CrossTask, COIN, UCF101, and ActivityNet, covering goal inference and action recognition tasks under open-vocabulary settings without requiring any training or fine-tuning. The results show that VidTFS outperforms pretrained and instruction-tuned multimodal language models that directly stack LLM and VFM for downstream video inference tasks. Our VidTFS with its adaptability shows the future potential for generalizing to new training-free video inference tasks.
title Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.12471