Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahmed, Umer, Mahmood, Syed Ahmed, Fateh, Fawad Javed, Luqman, M. Shaheer, Zia, M. Zeeshan, Tran, Quoc-Huy
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910135904894976
author Ahmed, Umer
Mahmood, Syed Ahmed
Fateh, Fawad Javed
Luqman, M. Shaheer
Zia, M. Zeeshan
Tran, Quoc-Huy
author_facet Ahmed, Umer
Mahmood, Syed Ahmed
Fateh, Fawad Javed
Luqman, M. Shaheer
Zia, M. Zeeshan
Tran, Quoc-Huy
contents We propose a novel hierarchical spatiotemporal vector quantization framework for unsupervised skeleton-based temporal action segmentation. We first introduce a hierarchical approach, which includes two consecutive levels of vector quantization. Specifically, the lower level associates skeletons with fine-grained subactions, while the higher level further aggregates subactions into action-level representations. Our hierarchical approach outperforms the non-hierarchical baseline, while primarily exploiting spatial cues by reconstructing input skeletons. Next, we extend our approach by leveraging both spatial and temporal information, yielding a hierarchical spatiotemporal vector quantization scheme. In particular, our hierarchical spatiotemporal approach performs multi-level clustering, while simultaneously recovering input skeletons and their corresponding timestamps. Lastly, extensive experiments on multiple benchmarks, including HuGaDB, LARa, and BABEL, demonstrate that our approach establishes a new state-of-the-art performance and reduces segment length bias in unsupervised skeleton-based temporal action segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_15196
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization
Ahmed, Umer
Mahmood, Syed Ahmed
Fateh, Fawad Javed
Luqman, M. Shaheer
Zia, M. Zeeshan
Tran, Quoc-Huy
Computer Vision and Pattern Recognition
We propose a novel hierarchical spatiotemporal vector quantization framework for unsupervised skeleton-based temporal action segmentation. We first introduce a hierarchical approach, which includes two consecutive levels of vector quantization. Specifically, the lower level associates skeletons with fine-grained subactions, while the higher level further aggregates subactions into action-level representations. Our hierarchical approach outperforms the non-hierarchical baseline, while primarily exploiting spatial cues by reconstructing input skeletons. Next, we extend our approach by leveraging both spatial and temporal information, yielding a hierarchical spatiotemporal vector quantization scheme. In particular, our hierarchical spatiotemporal approach performs multi-level clustering, while simultaneously recovering input skeletons and their corresponding timestamps. Lastly, extensive experiments on multiple benchmarks, including HuGaDB, LARa, and BABEL, demonstrate that our approach establishes a new state-of-the-art performance and reduces segment length bias in unsupervised skeleton-based temporal action segmentation.
title Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.15196