EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Flynn, John, Paier, Wolfgang, Dinev, Dimitar, Nguyen, Sam Nhut, Poghosyan, Hayk, Toribio, Manuel, Banerjee, Sandipan, Gafni, Guy
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917232946184192
author Flynn, John
Paier, Wolfgang
Dinev, Dimitar
Nguyen, Sam Nhut
Poghosyan, Hayk
Toribio, Manuel
Banerjee, Sandipan
Gafni, Guy
author_facet Flynn, John
Paier, Wolfgang
Dinev, Dimitar
Nguyen, Sam Nhut
Poghosyan, Hayk
Toribio, Manuel
Banerjee, Sandipan
Gafni, Guy
contents Current generative video models excel at producing novel content from text and image prompts, but leave a critical gap in editing existing pre-recorded videos, where minor alterations to the spoken script require preserving motion, temporal coherence, speaker identity, and accurate lip synchronization. We introduce EditYourself, a DiT-based framework for audio-driven video-to-video (V2V) editing that enables transcript-based modification of talking head videos, including the seamless addition, removal, and retiming of visually spoken content. Building on a general-purpose video diffusion model, EditYourself augments its V2V capabilities with audio conditioning and region-aware, edit-focused training extensions. This enables precise lip synchronization and temporally coherent restructuring of existing performances via spatiotemporal inpainting, including the synthesis of realistic human motion in newly added segments, while maintaining visual fidelity and identity consistency over long durations. This work represents a foundational step toward generative video models as practical tools for professional video post-production.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22127
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
Flynn, John
Paier, Wolfgang
Dinev, Dimitar
Nguyen, Sam Nhut
Poghosyan, Hayk
Toribio, Manuel
Banerjee, Sandipan
Gafni, Guy
Computer Vision and Pattern Recognition
Graphics
Machine Learning
Multimedia
Current generative video models excel at producing novel content from text and image prompts, but leave a critical gap in editing existing pre-recorded videos, where minor alterations to the spoken script require preserving motion, temporal coherence, speaker identity, and accurate lip synchronization. We introduce EditYourself, a DiT-based framework for audio-driven video-to-video (V2V) editing that enables transcript-based modification of talking head videos, including the seamless addition, removal, and retiming of visually spoken content. Building on a general-purpose video diffusion model, EditYourself augments its V2V capabilities with audio conditioning and region-aware, edit-focused training extensions. This enables precise lip synchronization and temporally coherent restructuring of existing performances via spatiotemporal inpainting, including the synthesis of realistic human motion in newly added segments, while maintaining visual fidelity and identity consistency over long durations. This work represents a foundational step toward generative video models as practical tools for professional video post-production.
title EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
Multimedia
url https://arxiv.org/abs/2601.22127