Pose-Aware Weakly-Supervised Action Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Seth Z., Ghoddoosian, Reza, Dwivedi, Isht, Agarwal, Nakul, Dariush, Behzad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909571359965184
author Zhao, Seth Z.
Ghoddoosian, Reza
Dwivedi, Isht
Agarwal, Nakul
Dariush, Behzad
author_facet Zhao, Seth Z.
Ghoddoosian, Reza
Dwivedi, Isht
Agarwal, Nakul
Dariush, Behzad
contents Understanding human behavior is an important problem in the pursuit of visual intelligence. A challenge in this endeavor is the extensive and costly effort required to accurately label action segments. To address this issue, we consider learning methods that demand minimal supervision for segmentation of human actions in long instructional videos. Specifically, we introduce a weakly-supervised framework that uniquely incorporates pose knowledge during training while omitting its use during inference, thereby distilling pose knowledge pertinent to each action component. We propose a pose-inspired contrastive loss as a part of the whole weakly-supervised framework which is trained to distinguish action boundaries more effectively. Our approach, validated through extensive experiments on representative datasets, outperforms previous state-of-the-art (SOTA) in segmenting long instructional videos under both online and offline settings. Additionally, we demonstrate the framework's adaptability to various segmentation backbones and pose extractors across different datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05700
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pose-Aware Weakly-Supervised Action Segmentation
Zhao, Seth Z.
Ghoddoosian, Reza
Dwivedi, Isht
Agarwal, Nakul
Dariush, Behzad
Computer Vision and Pattern Recognition
Understanding human behavior is an important problem in the pursuit of visual intelligence. A challenge in this endeavor is the extensive and costly effort required to accurately label action segments. To address this issue, we consider learning methods that demand minimal supervision for segmentation of human actions in long instructional videos. Specifically, we introduce a weakly-supervised framework that uniquely incorporates pose knowledge during training while omitting its use during inference, thereby distilling pose knowledge pertinent to each action component. We propose a pose-inspired contrastive loss as a part of the whole weakly-supervised framework which is trained to distinguish action boundaries more effectively. Our approach, validated through extensive experiments on representative datasets, outperforms previous state-of-the-art (SOTA) in segmenting long instructional videos under both online and offline settings. Additionally, we demonstrate the framework's adaptability to various segmentation backbones and pose extractors across different datasets.
title Pose-Aware Weakly-Supervised Action Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.05700