Deep Understanding of Sign Language for Sign to Subtitle Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jang, Youngjoon, Choi, Jeongsoo, Ahn, Junseok, Chung, Joon Son
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910859515658240
author Jang, Youngjoon
Choi, Jeongsoo
Ahn, Junseok
Chung, Joon Son
author_facet Jang, Youngjoon
Choi, Jeongsoo
Ahn, Junseok
Chung, Joon Son
contents The objective of this work is to align asynchronous subtitles in sign language videos with limited labelled data. To achieve this goal, we propose a novel framework with the following contributions: (1) we leverage fundamental grammatical rules of British Sign Language (BSL) to pre-process the input subtitles, (2) we design a selective alignment loss to optimise the model for predicting the temporal location of signs only when the queried sign actually occurs in a scene, and (3) we conduct self-training with refined pseudo-labels which are more accurate than the heuristic audio-aligned labels. From this, our model not only better understands the correlation between the text and the signs, but also holds potential for application in the translation of sign languages, particularly in scenarios where manual labelling of large-scale sign data is impractical or challenging. Extensive experimental results demonstrate that our approach achieves state-of-the-art results, surpassing previous baselines by substantial margins in terms of both frame-level accuracy and F1-score. This highlights the effectiveness and practicality of our framework in advancing the field of sign language video alignment and translation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03287
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Understanding of Sign Language for Sign to Subtitle Alignment
Jang, Youngjoon
Choi, Jeongsoo
Ahn, Junseok
Chung, Joon Son
Computer Vision and Pattern Recognition
The objective of this work is to align asynchronous subtitles in sign language videos with limited labelled data. To achieve this goal, we propose a novel framework with the following contributions: (1) we leverage fundamental grammatical rules of British Sign Language (BSL) to pre-process the input subtitles, (2) we design a selective alignment loss to optimise the model for predicting the temporal location of signs only when the queried sign actually occurs in a scene, and (3) we conduct self-training with refined pseudo-labels which are more accurate than the heuristic audio-aligned labels. From this, our model not only better understands the correlation between the text and the signs, but also holds potential for application in the translation of sign languages, particularly in scenarios where manual labelling of large-scale sign data is impractical or challenging. Extensive experimental results demonstrate that our approach achieves state-of-the-art results, surpassing previous baselines by substantial margins in terms of both frame-level accuracy and F1-score. This highlights the effectiveness and practicality of our framework in advancing the field of sign language video alignment and translation.
title Deep Understanding of Sign Language for Sign to Subtitle Alignment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.03287