MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Sung-Bin, Kim, Chae-Yeon, Lee, Son, Gihun, Hyun-Bin, Oh, Ju, Janghoon, Nam, Suekyeong, Oh, Tae-Hyun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation Metrics
by: Chae-Yeon, Lee, et al.
Published: (2025)
by: Chae-Yeon, Lee, et al.
Published: (2025)
Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading Expert
by: EunGi, Han, et al.
Published: (2024)
by: EunGi, Han, et al.
Published: (2024)
FPGS: Feed-Forward Semantic-aware Photorealistic Style Transfer of Large-Scale Gaussian Splatting
by: Kim, GeonU, et al.
Published: (2025)
by: Kim, GeonU, et al.
Published: (2025)
FPRF: Feed-Forward Photorealistic Style Transfer of Large-Scale 3D Neural Radiance Fields
by: Kim, GeonU, et al.
Published: (2024)
by: Kim, GeonU, et al.
Published: (2024)
OT-Talk: Animating 3D Talking Head with Optimal Transportation
by: Wang, Xinmu, et al.
Published: (2025)
by: Wang, Xinmu, et al.
Published: (2025)
Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering
by: Youwang, Kim, et al.
Published: (2023)
by: Youwang, Kim, et al.
Published: (2023)
FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling
by: Sung-Bin, Kim, et al.
Published: (2025)
by: Sung-Bin, Kim, et al.
Published: (2025)
MoDA: Multi-modal Diffusion Architecture for Talking Head Generation
by: Li, Xinyang, et al.
Published: (2025)
by: Li, Xinyang, et al.
Published: (2025)
StyGazeTalk: Learning Stylized Generation of Gaze and Head Dynamics
by: Shi, Chengwei, et al.
Published: (2025)
by: Shi, Chengwei, et al.
Published: (2025)
Supervising 3D Talking Head Avatars with Analysis-by-Audio-Synthesis
by: Daněček, Radek, et al.
Published: (2025)
by: Daněček, Radek, et al.
Published: (2025)
SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models
by: Hyun, Lee, et al.
Published: (2023)
by: Hyun, Lee, et al.
Published: (2023)
MoNeRF: Deformable Neural Rendering for Talking Heads via Latent Motion Navigation
by: X. Li, et al.
Published: (2024)
by: X. Li, et al.
Published: (2024)
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
by: Flynn, John, et al.
Published: (2026)
by: Flynn, John, et al.
Published: (2026)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
by: Xie, Yifan, et al.
Published: (2024)
by: Xie, Yifan, et al.
Published: (2024)
Learn2Talk: 3D Talking Face Learns from 2D Talking Face
by: Zhuang, Yixiang, et al.
Published: (2024)
by: Zhuang, Yixiang, et al.
Published: (2024)
PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis
by: Huang, Yihuan, et al.
Published: (2025)
by: Huang, Yihuan, et al.
Published: (2025)
One Shot, One Talk: Whole-body Talking Avatar from a Single Image
by: Xiang, Jun, et al.
Published: (2024)
by: Xiang, Jun, et al.
Published: (2024)
Learning Correlation-aware Aleatoric Uncertainty for 3D Hand Pose Estimation
by: Chae-Yeon, Lee, et al.
Published: (2025)
by: Chae-Yeon, Lee, et al.
Published: (2025)
Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions
by: Rakesh, Vineet Kumar, et al.
Published: (2025)
by: Rakesh, Vineet Kumar, et al.
Published: (2025)
VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
by: Cha, SeungJu, et al.
Published: (2025)
by: Cha, SeungJu, et al.
Published: (2025)
Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation
by: Long, Jianzhi, et al.
Published: (2025)
by: Long, Jianzhi, et al.
Published: (2025)
A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation
by: Airale, Louis, et al.
Published: (2023)
by: Airale, Louis, et al.
Published: (2023)
THGS: Lifelike Talking Human Avatar Synthesis From Monocular Video Via 3D Gaussian Splatting
by: Chuang Chen, et al.
Published: (2025)
by: Chuang Chen, et al.
Published: (2025)
Spatially and Temporally Optimized Audio‐Driven Talking Face Generation
by: Biao Dong, et al.
Published: (2024)
by: Biao Dong, et al.
Published: (2024)
The Life and Legacy of Bui Tuong Phong
by: Oh, Yoehan, et al.
Published: (2024)
by: Oh, Yoehan, et al.
Published: (2024)
TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
by: Low, Chetwin, et al.
Published: (2025)
by: Low, Chetwin, et al.
Published: (2025)
DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion Models
by: Sun, Zhiyao, et al.
Published: (2023)
by: Sun, Zhiyao, et al.
Published: (2023)
READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
by: Wang, Haotian, et al.
Published: (2025)
by: Wang, Haotian, et al.
Published: (2025)
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
by: Sung-Bin, Kim, et al.
Published: (2024)
by: Sung-Bin, Kim, et al.
Published: (2024)
FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models
by: Aneja, Shivangi, et al.
Published: (2023)
by: Aneja, Shivangi, et al.
Published: (2023)
Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head Synthesis
by: Shen, Shuai, et al.
Published: (2025)
by: Shen, Shuai, et al.
Published: (2025)
TetraSDF: Precise Mesh Extraction with Multi-resolution Tetrahedral Grid
by: Oh, Seonghun, et al.
Published: (2025)
by: Oh, Seonghun, et al.
Published: (2025)
PAColorHolo: A Perceptually-Aware Color Management Framework for Holographic Displays
by: Chen, Chun, et al.
Published: (2026)
by: Chen, Chun, et al.
Published: (2026)
Text-driven Talking Face Synthesis by Reprogramming Audio-driven Models
by: Choi, Jeongsoo, et al.
Published: (2023)
by: Choi, Jeongsoo, et al.
Published: (2023)
Generative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D Reconstruction
by: Nam, Seungtae, et al.
Published: (2024)
by: Nam, Seungtae, et al.
Published: (2024)
DC-VSR: Spatially and Temporally Consistent Video Super-Resolution with Video Diffusion Prior
by: Han, Janghyeok, et al.
Published: (2025)
by: Han, Janghyeok, et al.
Published: (2025)
PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
by: Zhang, Longhao, et al.
Published: (2024)
by: Zhang, Longhao, et al.
Published: (2024)
PhysTalk: Language-driven Real-time Physics in 3D Gaussian Scenes
by: Collorone, Luca, et al.
Published: (2025)
by: Collorone, Luca, et al.
Published: (2025)
Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation
by: Gan, Yuan, et al.
Published: (2025)
by: Gan, Yuan, et al.
Published: (2025)
MultiTalk: Introspective and Extrospective Dialogue for Human-Environment-LLM Alignment
by: Devarakonda, Venkata Naren, et al.
Published: (2024)
by: Devarakonda, Venkata Naren, et al.
Published: (2024)
Similar Items
-
Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation Metrics
by: Chae-Yeon, Lee, et al.
Published: (2025) -
Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading Expert
by: EunGi, Han, et al.
Published: (2024) -
FPGS: Feed-Forward Semantic-aware Photorealistic Style Transfer of Large-Scale Gaussian Splatting
by: Kim, GeonU, et al.
Published: (2025) -
FPRF: Feed-Forward Photorealistic Style Transfer of Large-Scale 3D Neural Radiance Fields
by: Kim, GeonU, et al.
Published: (2024) -
OT-Talk: Animating 3D Talking Head with Optimal Transportation
by: Wang, Xinmu, et al.
Published: (2025)