SkelCap: Automated Generation of Descriptive Text from Skeleton Keypoint Sequences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Keskin, Ali Emre, Keles, Hacer Yalim
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916332423872512
author Keskin, Ali Emre
Keles, Hacer Yalim
author_facet Keskin, Ali Emre
Keles, Hacer Yalim
contents Numerous sign language datasets exist, yet they typically cover only a limited selection of the thousands of signs used globally. Moreover, creating diverse sign language datasets is an expensive and challenging task due to the costs associated with gathering a varied group of signers. Motivated by these challenges, we aimed to develop a solution that addresses these limitations. In this context, we focused on textually describing body movements from skeleton keypoint sequences, leading to the creation of a new dataset. We structured this dataset around AUTSL, a comprehensive isolated Turkish sign language dataset. We also developed a baseline model, SkelCap, which can generate textual descriptions of body movements. This model processes the skeleton keypoints data as a vector, applies a fully connected layer for embedding, and utilizes a transformer neural network for sequence-to-sequence modeling. We conducted extensive evaluations of our model, including signer-agnostic and sign-agnostic assessments. The model achieved promising results, with a ROUGE-L score of 0.98 and a BLEU-4 score of 0.94 in the signer-agnostic evaluation. The dataset we have prepared, namely the AUTSL-SkelCap, will be made publicly available soon.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02977
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SkelCap: Automated Generation of Descriptive Text from Skeleton Keypoint Sequences
Keskin, Ali Emre
Keles, Hacer Yalim
Computer Vision and Pattern Recognition
Machine Learning
Numerous sign language datasets exist, yet they typically cover only a limited selection of the thousands of signs used globally. Moreover, creating diverse sign language datasets is an expensive and challenging task due to the costs associated with gathering a varied group of signers. Motivated by these challenges, we aimed to develop a solution that addresses these limitations. In this context, we focused on textually describing body movements from skeleton keypoint sequences, leading to the creation of a new dataset. We structured this dataset around AUTSL, a comprehensive isolated Turkish sign language dataset. We also developed a baseline model, SkelCap, which can generate textual descriptions of body movements. This model processes the skeleton keypoints data as a vector, applies a fully connected layer for embedding, and utilizes a transformer neural network for sequence-to-sequence modeling. We conducted extensive evaluations of our model, including signer-agnostic and sign-agnostic assessments. The model achieved promising results, with a ROUGE-L score of 0.98 and a BLEU-4 score of 0.94 in the signer-agnostic evaluation. The dataset we have prepared, namely the AUTSL-SkelCap, will be made publicly available soon.
title SkelCap: Automated Generation of Descriptive Text from Skeleton Keypoint Sequences
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2405.02977