Isolated Sign Language Recognition with Segmentation and Pose Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Perkins, Daniel, Hunter, Davis, Patel, Dhrumil, Flanagan, Galen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911323913191424
author Perkins, Daniel
Hunter, Davis
Patel, Dhrumil
Flanagan, Galen
author_facet Perkins, Daniel
Hunter, Davis
Patel, Dhrumil
Flanagan, Galen
contents The recent surge in large language models has automated translations of spoken and written languages. However, these advances remain largely inaccessible to American Sign Language (ASL) users, whose language relies on complex visual cues. Isolated sign language recognition (ISLR) - the task of classifying videos of individual signs - can help bridge this gap but is currently limited by scarce per-sign data, high signer variability, and substantial computational costs. We propose a model for ISLR that reduces computational requirements while maintaining robustness to signer variation. Our approach integrates (i) a pose estimation pipeline to extract hand and face joint coordinates, (ii) a segmentation module that isolates relevant information, and (iii) a ResNet-Transformer backbone to jointly model spatial and temporal dependencies.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14876
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Isolated Sign Language Recognition with Segmentation and Pose Estimation
Perkins, Daniel
Hunter, Davis
Patel, Dhrumil
Flanagan, Galen
Computer Vision and Pattern Recognition
68T07, 68T45, 68U10
The recent surge in large language models has automated translations of spoken and written languages. However, these advances remain largely inaccessible to American Sign Language (ASL) users, whose language relies on complex visual cues. Isolated sign language recognition (ISLR) - the task of classifying videos of individual signs - can help bridge this gap but is currently limited by scarce per-sign data, high signer variability, and substantial computational costs. We propose a model for ISLR that reduces computational requirements while maintaining robustness to signer variation. Our approach integrates (i) a pose estimation pipeline to extract hand and face joint coordinates, (ii) a segmentation module that isolates relevant information, and (iii) a ResNet-Transformer backbone to jointly model spatial and temporal dependencies.
title Isolated Sign Language Recognition with Segmentation and Pose Estimation
topic Computer Vision and Pattern Recognition
68T07, 68T45, 68U10
url https://arxiv.org/abs/2512.14876