EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Flynn, John, Paier, Wolfgang, Dinev, Dimitar, Nguyen, Sam Nhut, Poghosyan, Hayk, Toribio, Manuel, Banerjee, Sandipan, Gafni, Guy |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
by: Xie, Yifan, et al.
Published: (2024)
by: Xie, Yifan, et al.
Published: (2024)
Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation
by: Huang, Zikai, et al.
Published: (2025)
by: Huang, Zikai, et al.
Published: (2025)
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis
by: Liu, Xiaoxing, et al.
Published: (2025)
by: Liu, Xiaoxing, et al.
Published: (2025)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
by: Guan, Jiazhi, et al.
Published: (2025)
by: Guan, Jiazhi, et al.
Published: (2025)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
by: Hoe, Jiun Tian, et al.
Published: (2025)
by: Hoe, Jiun Tian, et al.
Published: (2025)
DASC: Depth-of-Field Aware Scene Complexity Metric for 3D Visualization on Light Field Display
by: Akbar, Kamran, et al.
Published: (2025)
by: Akbar, Kamran, et al.
Published: (2025)
GTLR-GS: Geometry-Texture Aware LiDAR-Regularized 3D Gaussian Splatting for Realistic Scene Reconstruction
by: Fang, Yan, et al.
Published: (2026)
by: Fang, Yan, et al.
Published: (2026)
SemanticGarment: Semantic-Controlled Generation and Editing of 3D Gaussian Garments
by: Wang, Ruiyan, et al.
Published: (2025)
by: Wang, Ruiyan, et al.
Published: (2025)
Generating Digital Models Using Text-to-3D and Image-to-3D Prompts: Critical Case Study
by: Ziatdinov, Rushan, et al.
Published: (2025)
by: Ziatdinov, Rushan, et al.
Published: (2025)
LAV: Audio-Driven Dynamic Visual Generation with Neural Compression and StyleGAN2
by: Jung, Jongmin, et al.
Published: (2025)
by: Jung, Jongmin, et al.
Published: (2025)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025)
by: Huang, Yiheng, et al.
Published: (2025)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
by: Guan, Jiazhi, et al.
Published: (2024)
by: Guan, Jiazhi, et al.
Published: (2024)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
SAGE: Semantic-Driven Adaptive Gaussian Splatting in Extended Reality
by: Schiavo, Chiara, et al.
Published: (2025)
by: Schiavo, Chiara, et al.
Published: (2025)
Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions
by: Rakesh, Vineet Kumar, et al.
Published: (2025)
by: Rakesh, Vineet Kumar, et al.
Published: (2025)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
by: Cheng, Shihao, et al.
Published: (2026)
by: Cheng, Shihao, et al.
Published: (2026)
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
by: Chen, Yi-Chun, et al.
Published: (2025)
by: Chen, Yi-Chun, et al.
Published: (2025)
On Copyright Risks of Text-to-Image Diffusion Models
by: Zhang, Yang, et al.
Published: (2023)
by: Zhang, Yang, et al.
Published: (2023)
Automatic Camera Trajectory Control with Enhanced Immersion for Virtual Cinematography
by: Wu, Xinyi, et al.
Published: (2023)
by: Wu, Xinyi, et al.
Published: (2023)
Real-Time Interactive Hybrid Ocean: Spectrum-Consistent Wave Particle-FFT Coupling
by: Xue, Shengze, et al.
Published: (2025)
by: Xue, Shengze, et al.
Published: (2025)
L3GS: Layered 3D Gaussian Splats for Efficient 3D Scene Delivery
by: Tsai, Yi-Zhen, et al.
Published: (2025)
by: Tsai, Yi-Zhen, et al.
Published: (2025)
From Code to Canvas
by: Werner, Bernhard O.
Published: (2025)
by: Werner, Bernhard O.
Published: (2025)
Do Inpainting Yourself: Generative Facial Inpainting Guided by Exemplars
by: Lu, Wanglong, et al.
Published: (2022)
by: Lu, Wanglong, et al.
Published: (2022)
Beyond the Desktop: XR-Driven Segmentation with Meta Quest 3 and MX Ink
by: de Paiva, Lisle Faray, et al.
Published: (2025)
by: de Paiva, Lisle Faray, et al.
Published: (2025)
Audio is all in one: speech-driven gesture synthetics using WavLM pre-trained model
by: Zhang, Fan, et al.
Published: (2023)
by: Zhang, Fan, et al.
Published: (2023)
Streaming Real-Time Rendered Scenes as 3D Gaussians
by: Siekkinen, Matti, et al.
Published: (2026)
by: Siekkinen, Matti, et al.
Published: (2026)
Towards Interactive Multimodal Representation of ML Functions for Human Understanding of ML
by: Wang, Bokang, et al.
Published: (2026)
by: Wang, Bokang, et al.
Published: (2026)
A Single Atlas is All You Need: Decoder-Side Gaussian Splatting for Immersive Video
by: Mieloch, Dawid, et al.
Published: (2026)
by: Mieloch, Dawid, et al.
Published: (2026)
Design of a UE5-based digital twin platform
by: Lyu, Shaoqiu, et al.
Published: (2024)
by: Lyu, Shaoqiu, et al.
Published: (2024)
Real-time 3D Light-field Viewing with Eye-tracking on Conventional Displays
by: Pham, Trung Hieu, et al.
Published: (2025)
by: Pham, Trung Hieu, et al.
Published: (2025)
"You'll Be Alice Adventuring in Wonderland!" Processes, Challenges, and Opportunities of Creating Animated Virtual Reality Stories
by: Yuan, Lin-Ping, et al.
Published: (2025)
by: Yuan, Lin-Ping, et al.
Published: (2025)
ScaleTrotter: Illustrative Visual Travels Across Negative Scales
by: Halladjian, Sarkis, et al.
Published: (2019)
by: Halladjian, Sarkis, et al.
Published: (2019)
Photoshop Batch Rendering Using Actions for Stylistic Video Editing
by: De La Fuente, Tessa
Published: (2025)
by: De La Fuente, Tessa
Published: (2025)
Resolution deficits drive simulator sickness and compromise reading performance in virtual environments
by: Wang, Jialin, et al.
Published: (2026)
by: Wang, Jialin, et al.
Published: (2026)
Coordinated 2D-3D Visualization of Volumetric Medical Data in XR with Multimodal Interactions
by: Liu, Qixuan, et al.
Published: (2025)
by: Liu, Qixuan, et al.
Published: (2025)
The perceptual gap between video see-through displays and natural human vision
by: Wang, Jialin, et al.
Published: (2026)
by: Wang, Jialin, et al.
Published: (2026)
XR is XR: Rethinking MR and XR as Neutral Umbrella Terms
by: Kurata, Takeshi
Published: (2026)
by: Kurata, Takeshi
Published: (2026)
Crafting Dynamic Virtual Activities with Advanced Multimodal Models
by: Li, Changyang, et al.
Published: (2024)
by: Li, Changyang, et al.
Published: (2024)
CvhSlicer 2.0: Immersive and Interactive Visualization of Chinese Visible Human Data in XR Environments
by: Qiu, Yue, et al.
Published: (2025)
by: Qiu, Yue, et al.
Published: (2025)
A Collaborative Extended Reality Prototype for 3D Surgical Planning and Visualization
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
Similar Items
-
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
by: Xie, Yifan, et al.
Published: (2024) -
Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation
by: Huang, Zikai, et al.
Published: (2025) -
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis
by: Liu, Xiaoxing, et al.
Published: (2025) -
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
by: Guan, Jiazhi, et al.
Published: (2025) -
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
by: Hoe, Jiun Tian, et al.
Published: (2025)