SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912461419970560 |
|---|---|
| author | Gueuwou, Shester Du, Xiaodan Shakhnarovich, Greg Livescu, Karen Liu, Alexander H. |
| author_facet | Gueuwou, Shester Du, Xiaodan Shakhnarovich, Greg Livescu, Karen Liu, Alexander H. |
| contents | Sign language processing has traditionally relied on task-specific models, limiting the potential for transfer learning across tasks. Pre-training methods for sign language have typically focused on either supervised pre-training, which cannot take advantage of unlabeled data, or context-independent (frame or video segment) representations, which ignore the effects of relationships across time in sign language. We introduce SHuBERT (Sign Hidden-Unit BERT), a self-supervised contextual representation model learned from approximately 1,000 hours of American Sign Language video. SHuBERT adapts masked token prediction objectives to multi-stream visual sign language input, learning to predict multiple targets corresponding to clustered hand, face, and body pose streams. SHuBERT achieves state-of-the-art performance across multiple tasks including sign language translation, isolated sign language recognition, and fingerspelling detection. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_16765 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction Gueuwou, Shester Du, Xiaodan Shakhnarovich, Greg Livescu, Karen Liu, Alexander H. Computation and Language Computer Vision and Pattern Recognition Sign language processing has traditionally relied on task-specific models, limiting the potential for transfer learning across tasks. Pre-training methods for sign language have typically focused on either supervised pre-training, which cannot take advantage of unlabeled data, or context-independent (frame or video segment) representations, which ignore the effects of relationships across time in sign language. We introduce SHuBERT (Sign Hidden-Unit BERT), a self-supervised contextual representation model learned from approximately 1,000 hours of American Sign Language video. SHuBERT adapts masked token prediction objectives to multi-stream visual sign language input, learning to predict multiple targets corresponding to clustered hand, face, and body pose streams. SHuBERT achieves state-of-the-art performance across multiple tasks including sign language translation, isolated sign language recognition, and fingerspelling detection. |
| title | SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction |
| topic | Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2411.16765 |