Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910135904894976 |
|---|---|
| author | Ahmed, Umer Mahmood, Syed Ahmed Fateh, Fawad Javed Luqman, M. Shaheer Zia, M. Zeeshan Tran, Quoc-Huy |
| author_facet | Ahmed, Umer Mahmood, Syed Ahmed Fateh, Fawad Javed Luqman, M. Shaheer Zia, M. Zeeshan Tran, Quoc-Huy |
| contents | We propose a novel hierarchical spatiotemporal vector quantization framework for unsupervised skeleton-based temporal action segmentation. We first introduce a hierarchical approach, which includes two consecutive levels of vector quantization. Specifically, the lower level associates skeletons with fine-grained subactions, while the higher level further aggregates subactions into action-level representations. Our hierarchical approach outperforms the non-hierarchical baseline, while primarily exploiting spatial cues by reconstructing input skeletons. Next, we extend our approach by leveraging both spatial and temporal information, yielding a hierarchical spatiotemporal vector quantization scheme. In particular, our hierarchical spatiotemporal approach performs multi-level clustering, while simultaneously recovering input skeletons and their corresponding timestamps. Lastly, extensive experiments on multiple benchmarks, including HuGaDB, LARa, and BABEL, demonstrate that our approach establishes a new state-of-the-art performance and reduces segment length bias in unsupervised skeleton-based temporal action segmentation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_15196 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization Ahmed, Umer Mahmood, Syed Ahmed Fateh, Fawad Javed Luqman, M. Shaheer Zia, M. Zeeshan Tran, Quoc-Huy Computer Vision and Pattern Recognition We propose a novel hierarchical spatiotemporal vector quantization framework for unsupervised skeleton-based temporal action segmentation. We first introduce a hierarchical approach, which includes two consecutive levels of vector quantization. Specifically, the lower level associates skeletons with fine-grained subactions, while the higher level further aggregates subactions into action-level representations. Our hierarchical approach outperforms the non-hierarchical baseline, while primarily exploiting spatial cues by reconstructing input skeletons. Next, we extend our approach by leveraging both spatial and temporal information, yielding a hierarchical spatiotemporal vector quantization scheme. In particular, our hierarchical spatiotemporal approach performs multi-level clustering, while simultaneously recovering input skeletons and their corresponding timestamps. Lastly, extensive experiments on multiple benchmarks, including HuGaDB, LARa, and BABEL, demonstrate that our approach establishes a new state-of-the-art performance and reduces segment length bias in unsupervised skeleton-based temporal action segmentation. |
| title | Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.15196 |