SignX: Continuous Sign Recognition in Compact Pose-Rich Latent Space

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Sen, Feng, Yalin, Sui, Chunyu, Zhong, Hongbin, Zhang, Yanxin, Yi, Hongwei, Hu, Hezhen, Metaxas, Dimitris N.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913039438053376
author Fang, Sen
Feng, Yalin
Sui, Chunyu
Zhong, Hongbin
Zhang, Yanxin
Yi, Hongwei
Hu, Hezhen
Metaxas, Dimitris N.
author_facet Fang, Sen
Feng, Yalin
Sui, Chunyu
Zhong, Hongbin
Zhang, Yanxin
Yi, Hongwei
Hu, Hezhen
Metaxas, Dimitris N.
contents The complexity of Sign Language (SL) data processing brings many challenges. The current approach to recognition of SL signs aims to translate RGB sign language videos through pose information into Word-based ID Glosses, which serve to uniquely identify signs. This paper proposes SignX, a novel framework for continuous sign language recognition (SLR) in compact pose-rich latent space. First, we construct a unified latent representation that encodes heterogeneous pose formats (SMPLer-X, DWPose, Mediapipe, PrimeDepth, and Sapiens Segmentation) into a compact, information-dense space. Second, we train a ViT-based Video-to-Pose module to extract this latent representation directly from raw videos. Finally, we develop a temporal modeling and sequence refinement method that operates entirely in this latent space. This multi-stage design achieves end-to-end SLR while significantly reducing computational consumption. Experimental results demonstrate that SignX achieves SOTA accuracy on continuous SLR and Translation task, delivering nearly a 50-fold acceleration over pixel-space baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16315
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SignX: Continuous Sign Recognition in Compact Pose-Rich Latent Space
Fang, Sen
Feng, Yalin
Sui, Chunyu
Zhong, Hongbin
Zhang, Yanxin
Yi, Hongwei
Hu, Hezhen
Metaxas, Dimitris N.
Computer Vision and Pattern Recognition
Computation and Language
The complexity of Sign Language (SL) data processing brings many challenges. The current approach to recognition of SL signs aims to translate RGB sign language videos through pose information into Word-based ID Glosses, which serve to uniquely identify signs. This paper proposes SignX, a novel framework for continuous sign language recognition (SLR) in compact pose-rich latent space. First, we construct a unified latent representation that encodes heterogeneous pose formats (SMPLer-X, DWPose, Mediapipe, PrimeDepth, and Sapiens Segmentation) into a compact, information-dense space. Second, we train a ViT-based Video-to-Pose module to extract this latent representation directly from raw videos. Finally, we develop a temporal modeling and sequence refinement method that operates entirely in this latent space. This multi-stage design achieves end-to-end SLR while significantly reducing computational consumption. Experimental results demonstrate that SignX achieves SOTA accuracy on continuous SLR and Translation task, delivering nearly a 50-fold acceleration over pixel-space baselines.
title SignX: Continuous Sign Recognition in Compact Pose-Rich Latent Space
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2504.16315