Express4D: Expressive, Friendly, and Extensible 4D Facial Motion Generation Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aloni, Yaron, Shalev-Arkushin, Rotem, Shafir, Yonatan, Tevet, Guy, Fried, Ohad, Bermano, Amit Haim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908492811468800
author Aloni, Yaron
Shalev-Arkushin, Rotem
Shafir, Yonatan
Tevet, Guy
Fried, Ohad
Bermano, Amit Haim
author_facet Aloni, Yaron
Shalev-Arkushin, Rotem
Shafir, Yonatan
Tevet, Guy
Fried, Ohad
Bermano, Amit Haim
contents Dynamic facial expression generation from natural language is a crucial task in Computer Graphics, with applications in Animation, Virtual Avatars, and Human-Computer Interaction. However, current generative models suffer from datasets that are either speech-driven or limited to coarse emotion labels, lacking the nuanced, expressive descriptions needed for fine-grained control, and were captured using elaborate and expensive equipment. We hence present a new dataset of facial motion sequences featuring nuanced performances and semantic annotation. The data is easily collected using commodity equipment and LLM-generated natural language instructions, in the popular ARKit blendshape format. This provides riggable motion, rich with expressive performances and labels. We accordingly train two baseline models, and evaluate their performance for future benchmarking. Using our Express4D dataset, the trained models can learn meaningful text-to-expression motion generation and capture the many-to-many mapping of the two modalities. The dataset, code, and video examples are available on our webpage: https://jaron1990.github.io/Express4D/
format Preprint
id arxiv_https___arxiv_org_abs_2508_12438
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Express4D: Expressive, Friendly, and Extensible 4D Facial Motion Generation Benchmark
Aloni, Yaron
Shalev-Arkushin, Rotem
Shafir, Yonatan
Tevet, Guy
Fried, Ohad
Bermano, Amit Haim
Graphics
Computer Vision and Pattern Recognition
Dynamic facial expression generation from natural language is a crucial task in Computer Graphics, with applications in Animation, Virtual Avatars, and Human-Computer Interaction. However, current generative models suffer from datasets that are either speech-driven or limited to coarse emotion labels, lacking the nuanced, expressive descriptions needed for fine-grained control, and were captured using elaborate and expensive equipment. We hence present a new dataset of facial motion sequences featuring nuanced performances and semantic annotation. The data is easily collected using commodity equipment and LLM-generated natural language instructions, in the popular ARKit blendshape format. This provides riggable motion, rich with expressive performances and labels. We accordingly train two baseline models, and evaluate their performance for future benchmarking. Using our Express4D dataset, the trained models can learn meaningful text-to-expression motion generation and capture the many-to-many mapping of the two modalities. The dataset, code, and video examples are available on our webpage: https://jaron1990.github.io/Express4D/
title Express4D: Expressive, Friendly, and Extensible 4D Facial Motion Generation Benchmark
topic Graphics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.12438