B2F: End-to-End Body-to-Face Motion Generation with Style Reference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jang, Bokyung, Jung, Eunho, Lee, Yoonsang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911271631192064
author Jang, Bokyung
Jung, Eunho
Lee, Yoonsang
author_facet Jang, Bokyung
Jung, Eunho
Lee, Yoonsang
contents Human motion naturally integrates body movements and facial expressions, forming a unified perception. If a virtual character's facial expression does not align well with its body movements, it may weaken the perception of the character as a cohesive whole. Motivated by this, we propose B2F, a model that generates facial motions aligned with body movements. B2F takes a facial style reference as input, generating facial animations that reflect the provided style while maintaining consistency with the associated body motion. To achieve this, B2F learns a disentangled representation of content and style, using alignment and consistency-based objectives. We represent style using discrete latent codes learned via the Gumbel-Softmax trick, enabling diverse expression generation with a structured latent representation. B2F outputs facial motion in the FLAME format, making it compatible with SMPL-X characters, and supports ARKit-style avatars through a dedicated conversion module. Our evaluations show that B2F generates expressive and engaging facial animations that synchronize with body movements and style intent, while mitigating perceptual dissonance from mismatched cues, and generalizing across diverse characters and styles.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13988
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle B2F: End-to-End Body-to-Face Motion Generation with Style Reference
Jang, Bokyung
Jung, Eunho
Lee, Yoonsang
Graphics
Human motion naturally integrates body movements and facial expressions, forming a unified perception. If a virtual character's facial expression does not align well with its body movements, it may weaken the perception of the character as a cohesive whole. Motivated by this, we propose B2F, a model that generates facial motions aligned with body movements. B2F takes a facial style reference as input, generating facial animations that reflect the provided style while maintaining consistency with the associated body motion. To achieve this, B2F learns a disentangled representation of content and style, using alignment and consistency-based objectives. We represent style using discrete latent codes learned via the Gumbel-Softmax trick, enabling diverse expression generation with a structured latent representation. B2F outputs facial motion in the FLAME format, making it compatible with SMPL-X characters, and supports ARKit-style avatars through a dedicated conversion module. Our evaluations show that B2F generates expressive and engaging facial animations that synchronize with body movements and style intent, while mitigating perceptual dissonance from mismatched cues, and generalizing across diverse characters and styles.
title B2F: End-to-End Body-to-Face Motion Generation with Style Reference
topic Graphics
url https://arxiv.org/abs/2511.13988