Uni-Sign: Toward Unified Sign Language Understanding at Scale

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Zecheng, Zhou, Wengang, Zhao, Weichao, Wu, Kepeng, Hu, Hezhen, Li, Houqiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910873573916672
author Li, Zecheng
Zhou, Wengang
Zhao, Weichao
Wu, Kepeng
Hu, Hezhen
Li, Houqiang
author_facet Li, Zecheng
Zhou, Wengang
Zhao, Weichao
Wu, Kepeng
Hu, Hezhen
Li, Houqiang
contents Sign language pre-training has gained increasing attention for its ability to enhance performance across various sign language understanding (SLU) tasks. However, existing methods often suffer from a gap between pre-training and fine-tuning, leading to suboptimal results. To address this, we propose Uni-Sign, a unified pre-training framework that eliminates the gap between pre-training and downstream SLU tasks through a large-scale generative pre-training strategy and a novel fine-tuning paradigm. First, we introduce CSL-News, a large-scale Chinese Sign Language (CSL) dataset containing 1,985 hours of video paired with textual annotations, which enables effective large-scale pre-training. Second, Uni-Sign unifies SLU tasks by treating downstream tasks as a single sign language translation (SLT) task during fine-tuning, ensuring seamless knowledge transfer between pre-training and fine-tuning. Furthermore, we incorporate a prior-guided fusion (PGF) module and a score-aware sampling strategy to efficiently fuse pose and RGB information, addressing keypoint inaccuracies and improving computational efficiency. Extensive experiments across multiple SLU benchmarks demonstrate that Uni-Sign achieves state-of-the-art performance across multiple downstream SLU tasks. Dataset and code are available at github.com/ZechengLi19/Uni-Sign.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15187
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uni-Sign: Toward Unified Sign Language Understanding at Scale
Li, Zecheng
Zhou, Wengang
Zhao, Weichao
Wu, Kepeng
Hu, Hezhen
Li, Houqiang
Computer Vision and Pattern Recognition
Sign language pre-training has gained increasing attention for its ability to enhance performance across various sign language understanding (SLU) tasks. However, existing methods often suffer from a gap between pre-training and fine-tuning, leading to suboptimal results. To address this, we propose Uni-Sign, a unified pre-training framework that eliminates the gap between pre-training and downstream SLU tasks through a large-scale generative pre-training strategy and a novel fine-tuning paradigm. First, we introduce CSL-News, a large-scale Chinese Sign Language (CSL) dataset containing 1,985 hours of video paired with textual annotations, which enables effective large-scale pre-training. Second, Uni-Sign unifies SLU tasks by treating downstream tasks as a single sign language translation (SLT) task during fine-tuning, ensuring seamless knowledge transfer between pre-training and fine-tuning. Furthermore, we incorporate a prior-guided fusion (PGF) module and a score-aware sampling strategy to efficiently fuse pose and RGB information, addressing keypoint inaccuracies and improving computational efficiency. Extensive experiments across multiple SLU benchmarks demonstrate that Uni-Sign achieves state-of-the-art performance across multiple downstream SLU tasks. Dataset and code are available at github.com/ZechengLi19/Uni-Sign.
title Uni-Sign: Toward Unified Sign Language Understanding at Scale
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.15187