CoFInAl: Enhancing Action Quality Assessment with Coarse-to-Fine Instruction Alignment

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Kanglei, Li, Junlin, Cai, Ruizhi, Wang, Liyuan, Zhang, Xingxing, Liang, Xiaohui
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912059391737856
author Zhou, Kanglei
Li, Junlin
Cai, Ruizhi
Wang, Liyuan
Zhang, Xingxing
Liang, Xiaohui
author_facet Zhou, Kanglei
Li, Junlin
Cai, Ruizhi
Wang, Liyuan
Zhang, Xingxing
Liang, Xiaohui
contents Action Quality Assessment (AQA) is pivotal for quantifying actions across domains like sports and medical care. Existing methods often rely on pre-trained backbones from large-scale action recognition datasets to boost performance on smaller AQA datasets. However, this common strategy yields suboptimal results due to the inherent struggle of these backbones to capture the subtle cues essential for AQA. Moreover, fine-tuning on smaller datasets risks overfitting. To address these issues, we propose Coarse-to-Fine Instruction Alignment (CoFInAl). Inspired by recent advances in large language model tuning, CoFInAl aligns AQA with broader pre-trained tasks by reformulating it as a coarse-to-fine classification task. Initially, it learns grade prototypes for coarse assessment and then utilizes fixed sub-grade prototypes for fine-grained assessment. This hierarchical approach mirrors the judging process, enhancing interpretability within the AQA framework. Experimental results on two long-term AQA datasets demonstrate CoFInAl achieves state-of-the-art performance with significant correlation gains of 5.49% and 3.55% on Rhythmic Gymnastics and Fis-V, respectively. Our code is available at https://github.com/ZhouKanglei/CoFInAl_AQA.
format Preprint
id arxiv_https___arxiv_org_abs_2404_13999
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CoFInAl: Enhancing Action Quality Assessment with Coarse-to-Fine Instruction Alignment
Zhou, Kanglei
Li, Junlin
Cai, Ruizhi
Wang, Liyuan
Zhang, Xingxing
Liang, Xiaohui
Computer Vision and Pattern Recognition
Action Quality Assessment (AQA) is pivotal for quantifying actions across domains like sports and medical care. Existing methods often rely on pre-trained backbones from large-scale action recognition datasets to boost performance on smaller AQA datasets. However, this common strategy yields suboptimal results due to the inherent struggle of these backbones to capture the subtle cues essential for AQA. Moreover, fine-tuning on smaller datasets risks overfitting. To address these issues, we propose Coarse-to-Fine Instruction Alignment (CoFInAl). Inspired by recent advances in large language model tuning, CoFInAl aligns AQA with broader pre-trained tasks by reformulating it as a coarse-to-fine classification task. Initially, it learns grade prototypes for coarse assessment and then utilizes fixed sub-grade prototypes for fine-grained assessment. This hierarchical approach mirrors the judging process, enhancing interpretability within the AQA framework. Experimental results on two long-term AQA datasets demonstrate CoFInAl achieves state-of-the-art performance with significant correlation gains of 5.49% and 3.55% on Rhythmic Gymnastics and Fis-V, respectively. Our code is available at https://github.com/ZhouKanglei/CoFInAl_AQA.
title CoFInAl: Enhancing Action Quality Assessment with Coarse-to-Fine Instruction Alignment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.13999