HOI-aware Adaptive Network for Weakly-supervised Action Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Runzhong, Wang, Suchen, Duan, Yueqi, Tang, Yansong, Zhang, Yue, Tan, Yap-Peng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910195103301632
author Zhang, Runzhong
Wang, Suchen
Duan, Yueqi
Tang, Yansong
Zhang, Yue
Tan, Yap-Peng
author_facet Zhang, Runzhong
Wang, Suchen
Duan, Yueqi
Tang, Yansong
Zhang, Yue
Tan, Yap-Peng
contents In this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed network to predict the action of each frame with the neighboring frames. However, this would result in ambiguity when estimating similar actions, such as pouring juice and pouring coffee. To address this, we aim to exploit temporally global but spatially local human-object interactions (HOI) as video-level prior knowledge for action segmentation. The long-term HOI sequence provides crucial contextual information to distinguish ambiguous actions, where our network dynamically adapts to the given HOI sequence at test time. More specifically, we first design a video HOI encoder that extracts, selects, and integrates the most representative HOI throughout the video. Then, we propose a two-branch HyperNetwork to learn an adaptive temporal encoder, which automatically adjusts the parameters based on the HOI information of various videos on the fly. Extensive experiments on two widely-used datasets including Breakfast and 50Salads demonstrate the effectiveness of our method under different evaluation metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26227
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HOI-aware Adaptive Network for Weakly-supervised Action Segmentation
Zhang, Runzhong
Wang, Suchen
Duan, Yueqi
Tang, Yansong
Zhang, Yue
Tan, Yap-Peng
Computer Vision and Pattern Recognition
In this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed network to predict the action of each frame with the neighboring frames. However, this would result in ambiguity when estimating similar actions, such as pouring juice and pouring coffee. To address this, we aim to exploit temporally global but spatially local human-object interactions (HOI) as video-level prior knowledge for action segmentation. The long-term HOI sequence provides crucial contextual information to distinguish ambiguous actions, where our network dynamically adapts to the given HOI sequence at test time. More specifically, we first design a video HOI encoder that extracts, selects, and integrates the most representative HOI throughout the video. Then, we propose a two-branch HyperNetwork to learn an adaptive temporal encoder, which automatically adjusts the parameters based on the HOI information of various videos on the fly. Extensive experiments on two widely-used datasets including Breakfast and 50Salads demonstrate the effectiveness of our method under different evaluation metrics.
title HOI-aware Adaptive Network for Weakly-supervised Action Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.26227