Exploring Timeline Control for Facial Motion Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Yifeng, Qi, Jinwei, Ji, Chaonan, Zhang, Peng, Zhang, Bang, Deng, Zhidong, Bo, Liefeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916761716129792
author Ma, Yifeng
Qi, Jinwei
Ji, Chaonan
Zhang, Peng
Zhang, Bang
Deng, Zhidong
Bo, Liefeng
author_facet Ma, Yifeng
Qi, Jinwei
Ji, Chaonan
Zhang, Peng
Zhang, Bang
Deng, Zhidong
Bo, Liefeng
contents This paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing. Users can specify a multi-track timeline of facial actions arranged in temporal intervals, allowing precise control over the timing of each action. To model the timeline control capability, We first annotate the time intervals of facial actions in natural facial motion sequences at a frame-level granularity. This process is facilitated by Toeplitz Inverse Covariance-based Clustering to minimize human labor. Based on the annotations, we propose a diffusion-based generation model capable of generating facial motions that are natural and accurately aligned with input timelines. Our method supports text-guided motion generation by using ChatGPT to convert text into timelines. Experimental results show that our method can annotate facial action intervals with satisfactory accuracy, and produces natural facial motions accurately aligned with timelines.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Timeline Control for Facial Motion Generation
Ma, Yifeng
Qi, Jinwei
Ji, Chaonan
Zhang, Peng
Zhang, Bang
Deng, Zhidong
Bo, Liefeng
Computer Vision and Pattern Recognition
This paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing. Users can specify a multi-track timeline of facial actions arranged in temporal intervals, allowing precise control over the timing of each action. To model the timeline control capability, We first annotate the time intervals of facial actions in natural facial motion sequences at a frame-level granularity. This process is facilitated by Toeplitz Inverse Covariance-based Clustering to minimize human labor. Based on the annotations, we propose a diffusion-based generation model capable of generating facial motions that are natural and accurately aligned with input timelines. Our method supports text-guided motion generation by using ChatGPT to convert text into timelines. Experimental results show that our method can annotate facial action intervals with satisfactory accuracy, and produces natural facial motions accurately aligned with timelines.
title Exploring Timeline Control for Facial Motion Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20861