Autoregressive Sign Language Production: A Gloss-Free Approach with Discrete Representations
Fuente:
arXiv
Guardado en:
| Autores principales: | , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866929377732722688 |
|---|---|
| author | Hwang, Eui Jun Lee, Huije Park, Jong C. |
| author_facet | Hwang, Eui Jun Lee, Huije Park, Jong C. |
| contents | Gloss-free Sign Language Production (SLP) offers a direct translation of spoken language sentences into sign language, bypassing the need for gloss intermediaries. This paper presents the Sign language Vector Quantization Network, a novel approach to SLP that leverages Vector Quantization to derive discrete representations from sign pose sequences. Our method, rooted in both manual and non-manual elements of signing, supports advanced decoding methods and integrates latent-level alignment for enhanced linguistic coherence. Through comprehensive evaluations, we demonstrate superior performance of our method over prior SLP methods and highlight the reliability of Back-Translation and Fréchet Gesture Distance as evaluation metrics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2309_12179 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Autoregressive Sign Language Production: A Gloss-Free Approach with Discrete Representations Hwang, Eui Jun Lee, Huije Park, Jong C. Computer Vision and Pattern Recognition Gloss-free Sign Language Production (SLP) offers a direct translation of spoken language sentences into sign language, bypassing the need for gloss intermediaries. This paper presents the Sign language Vector Quantization Network, a novel approach to SLP that leverages Vector Quantization to derive discrete representations from sign pose sequences. Our method, rooted in both manual and non-manual elements of signing, supports advanced decoding methods and integrates latent-level alignment for enhanced linguistic coherence. Through comprehensive evaluations, we demonstrate superior performance of our method over prior SLP methods and highlight the reliability of Back-Translation and Fréchet Gesture Distance as evaluation metrics. |
| title | Autoregressive Sign Language Production: A Gloss-Free Approach with Discrete Representations |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2309.12179 |