Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Jiahao, Chen, Yan, Xu, Mingwang, Shang, Hanlin, Chen, Yuxuan, Zhan, Yun, Dong, Zilong, Yao, Yao, Wang, Jingdong, Zhu, Siyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915644318941184
author Cui, Jiahao
Chen, Yan
Xu, Mingwang
Shang, Hanlin
Chen, Yuxuan
Zhan, Yun
Dong, Zilong
Yao, Yao
Wang, Jingdong
Zhu, Siyu
author_facet Cui, Jiahao
Chen, Yan
Xu, Mingwang
Shang, Hanlin
Chen, Yuxuan
Zhan, Yun
Dong, Zilong
Yao, Yao
Wang, Jingdong
Zhu, Siyu
contents Generating highly dynamic and photorealistic portrait animations driven by audio and skeletal motion remains challenging due to the need for precise lip synchronization, natural facial expressions, and high-fidelity body motion dynamics. We propose a human-preference-aligned diffusion framework that addresses these challenges through two key innovations. First, we introduce direct preference optimization tailored for human-centric animation, leveraging a curated dataset of human preferences to align generated outputs with perceptual metrics for portrait motion-video alignment and naturalness of expression. Second, the proposed temporal motion modulation resolves spatiotemporal resolution mismatches by reshaping motion conditions into dimensionally aligned latent features through temporal channel redistribution and proportional feature expansion, preserving the fidelity of high-frequency motion details in diffusion-based synthesis. The proposed mechanism is complementary to existing UNet and DiT-based portrait diffusion approaches, and experiments demonstrate obvious improvements in lip-audio synchronization, expression vividness, body motion coherence over baseline methods, alongside notable gains in human preference metrics. Our model and source code can be found at: https://github.com/fudan-generative-vision/hallo4.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23525
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization
Cui, Jiahao
Chen, Yan
Xu, Mingwang
Shang, Hanlin
Chen, Yuxuan
Zhan, Yun
Dong, Zilong
Yao, Yao
Wang, Jingdong
Zhu, Siyu
Computer Vision and Pattern Recognition
Generating highly dynamic and photorealistic portrait animations driven by audio and skeletal motion remains challenging due to the need for precise lip synchronization, natural facial expressions, and high-fidelity body motion dynamics. We propose a human-preference-aligned diffusion framework that addresses these challenges through two key innovations. First, we introduce direct preference optimization tailored for human-centric animation, leveraging a curated dataset of human preferences to align generated outputs with perceptual metrics for portrait motion-video alignment and naturalness of expression. Second, the proposed temporal motion modulation resolves spatiotemporal resolution mismatches by reshaping motion conditions into dimensionally aligned latent features through temporal channel redistribution and proportional feature expansion, preserving the fidelity of high-frequency motion details in diffusion-based synthesis. The proposed mechanism is complementary to existing UNet and DiT-based portrait diffusion approaches, and experiments demonstrate obvious improvements in lip-audio synchronization, expression vividness, body motion coherence over baseline methods, alongside notable gains in human preference metrics. Our model and source code can be found at: https://github.com/fudan-generative-vision/hallo4.
title Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.23525