Improving Temporal Action Segmentation via Constraint-Aware Decoding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ee, Yeo Keat, Roy, Debaditya, Li, Chen, Zhang, Hao, Fernando, Basura
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917479882686464
author Ee, Yeo Keat
Roy, Debaditya
Li, Chen
Zhang, Hao
Fernando, Basura
author_facet Ee, Yeo Keat
Roy, Debaditya
Li, Chen
Zhang, Hao
Fernando, Basura
contents Temporal action segmentation (TAS) divides untrimmed videos into labeled action segments. While fully supervised methods have advanced the field, challenges such as action variability, ambiguous boundaries, and high annotation costs remain, especially in new or low-resource domains. Grammar-based approaches improve segmentation with structural priors but rely on complex parsing limiting scalability. In this work, we propose a lightweight, constraint-based refinement framework that enhances TAS predictions by integrating statistical structural priors such as transition confidence, action boundary sets, and per-class duration, that can be directly extracted from annotated data. These constraints are integrated into a modified Viterbi decoding algorithm, allowing inference-time refinement without retraining or added model complexity. Our approach improves both fully and semi-supervised TAS models by correcting structural prediction errors while maintaining high efficiency. Code is available at https://github.com/LUNAProject22/CAD
format Preprint
id arxiv_https___arxiv_org_abs_2605_10149
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Improving Temporal Action Segmentation via Constraint-Aware Decoding
Ee, Yeo Keat
Roy, Debaditya
Li, Chen
Zhang, Hao
Fernando, Basura
Computer Vision and Pattern Recognition
Temporal action segmentation (TAS) divides untrimmed videos into labeled action segments. While fully supervised methods have advanced the field, challenges such as action variability, ambiguous boundaries, and high annotation costs remain, especially in new or low-resource domains. Grammar-based approaches improve segmentation with structural priors but rely on complex parsing limiting scalability. In this work, we propose a lightweight, constraint-based refinement framework that enhances TAS predictions by integrating statistical structural priors such as transition confidence, action boundary sets, and per-class duration, that can be directly extracted from annotated data. These constraints are integrated into a modified Viterbi decoding algorithm, allowing inference-time refinement without retraining or added model complexity. Our approach improves both fully and semi-supervised TAS models by correcting structural prediction errors while maintaining high efficiency. Code is available at https://github.com/LUNAProject22/CAD
title Improving Temporal Action Segmentation via Constraint-Aware Decoding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.10149