Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions
Fuente:
arXiv
Saved in:
| Main Authors: | Rakesh, Vineet Kumar, Mazumdar, Soumya, Maity, Research Pratim, Pal, Sarbajit, Das, Amitabha, Samanta, Tapas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction
by: Rakesh, Vineet Kumar, et al.
Published: (2026)
by: Rakesh, Vineet Kumar, et al.
Published: (2026)
BayesFusion-SDF: Probabilistic Signed Distance Fusion with View Planning on CPU
by: Mazumdar, Soumya, et al.
Published: (2026)
by: Mazumdar, Soumya, et al.
Published: (2026)
VedicTHG: Symbolic Vedic Computation for Low-Resource Talking-Head Generation in Educational Avatars
by: Rakesh, Vineet Kumar, et al.
Published: (2026)
by: Rakesh, Vineet Kumar, et al.
Published: (2026)
PrivFedTalk: Privacy-Aware Federated Diffusion with Identity-Stable Adapters for Personalized Talking-Head Generation
by: Mazumdar, Soumya, et al.
Published: (2026)
by: Mazumdar, Soumya, et al.
Published: (2026)
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
by: Flynn, John, et al.
Published: (2026)
by: Flynn, John, et al.
Published: (2026)
Large Language Models for Computer-Aided Design: A Survey
by: Zhang, Licheng, et al.
Published: (2025)
by: Zhang, Licheng, et al.
Published: (2025)
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis
by: Liu, Xiaoxing, et al.
Published: (2025)
by: Liu, Xiaoxing, et al.
Published: (2025)
Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
by: Hu, Xiaowei, et al.
Published: (2024)
by: Hu, Xiaowei, et al.
Published: (2024)
Analysis of Hyperparameter Optimization Effects on Lightweight Deep Models for Real-Time Image Classification
by: Rakesh, Vineet Kumar, et al.
Published: (2025)
by: Rakesh, Vineet Kumar, et al.
Published: (2025)
MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
by: Gupta, Prerit, et al.
Published: (2025)
by: Gupta, Prerit, et al.
Published: (2025)
DASC: Depth-of-Field Aware Scene Complexity Metric for 3D Visualization on Light Field Display
by: Akbar, Kamran, et al.
Published: (2025)
by: Akbar, Kamran, et al.
Published: (2025)
Towards Interactive Multimodal Representation of ML Functions for Human Understanding of ML
by: Wang, Bokang, et al.
Published: (2026)
by: Wang, Bokang, et al.
Published: (2026)
Design of a UE5-based digital twin platform
by: Lyu, Shaoqiu, et al.
Published: (2024)
by: Lyu, Shaoqiu, et al.
Published: (2024)
Real-time 3D Light-field Viewing with Eye-tracking on Conventional Displays
by: Pham, Trung Hieu, et al.
Published: (2025)
by: Pham, Trung Hieu, et al.
Published: (2025)
"You'll Be Alice Adventuring in Wonderland!" Processes, Challenges, and Opportunities of Creating Animated Virtual Reality Stories
by: Yuan, Lin-Ping, et al.
Published: (2025)
by: Yuan, Lin-Ping, et al.
Published: (2025)
ScaleTrotter: Illustrative Visual Travels Across Negative Scales
by: Halladjian, Sarkis, et al.
Published: (2019)
by: Halladjian, Sarkis, et al.
Published: (2019)
Photoshop Batch Rendering Using Actions for Stylistic Video Editing
by: De La Fuente, Tessa
Published: (2025)
by: De La Fuente, Tessa
Published: (2025)
Resolution deficits drive simulator sickness and compromise reading performance in virtual environments
by: Wang, Jialin, et al.
Published: (2026)
by: Wang, Jialin, et al.
Published: (2026)
Coordinated 2D-3D Visualization of Volumetric Medical Data in XR with Multimodal Interactions
by: Liu, Qixuan, et al.
Published: (2025)
by: Liu, Qixuan, et al.
Published: (2025)
The perceptual gap between video see-through displays and natural human vision
by: Wang, Jialin, et al.
Published: (2026)
by: Wang, Jialin, et al.
Published: (2026)
XR is XR: Rethinking MR and XR as Neutral Umbrella Terms
by: Kurata, Takeshi
Published: (2026)
by: Kurata, Takeshi
Published: (2026)
Crafting Dynamic Virtual Activities with Advanced Multimodal Models
by: Li, Changyang, et al.
Published: (2024)
by: Li, Changyang, et al.
Published: (2024)
CvhSlicer 2.0: Immersive and Interactive Visualization of Chinese Visible Human Data in XR Environments
by: Qiu, Yue, et al.
Published: (2025)
by: Qiu, Yue, et al.
Published: (2025)
A Collaborative Extended Reality Prototype for 3D Surgical Planning and Visualization
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
A Survey on 3D Gaussian Splatting
by: Chen, Guikun, et al.
Published: (2024)
by: Chen, Guikun, et al.
Published: (2024)
Beyond the Desktop: XR-Driven Segmentation with Meta Quest 3 and MX Ink
by: de Paiva, Lisle Faray, et al.
Published: (2025)
by: de Paiva, Lisle Faray, et al.
Published: (2025)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
by: Xie, Yifan, et al.
Published: (2024)
by: Xie, Yifan, et al.
Published: (2024)
EyeNavGS: A 6-DoF Navigation Dataset and Record-n-Replay Software for Real-World 3DGS Scenes in VR
by: Ding, Zihao, et al.
Published: (2025)
by: Ding, Zihao, et al.
Published: (2025)
VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
by: Cha, SeungJu, et al.
Published: (2025)
by: Cha, SeungJu, et al.
Published: (2025)
Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
by: Lin, Jiantao, et al.
Published: (2025)
by: Lin, Jiantao, et al.
Published: (2025)
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
by: Dong, Wenqi, et al.
Published: (2025)
by: Dong, Wenqi, et al.
Published: (2025)
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
by: Zheng, Jiayi, et al.
Published: (2025)
by: Zheng, Jiayi, et al.
Published: (2025)
Perceive-Sample-Compress: Towards Real-Time 3D Gaussian Splatting
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
by: Chen, Yi-Chun, et al.
Published: (2025)
by: Chen, Yi-Chun, et al.
Published: (2025)
Splatography: Sparse multi-view dynamic Gaussian Splatting for filmmaking challenges
by: Azzarelli, Adrian, et al.
Published: (2025)
by: Azzarelli, Adrian, et al.
Published: (2025)
DreamCinema: Cinematic Transfer with Free Camera and 3D Character
by: Chen, Weiliang, et al.
Published: (2024)
by: Chen, Weiliang, et al.
Published: (2024)
MesonGS++: Post-training Compression of 3D Gaussian Splatting with Hyperparameter Searching
by: Xie, Shuzhao, et al.
Published: (2026)
by: Xie, Shuzhao, et al.
Published: (2026)
altiro3D: Scene representation from single image and novel view synthesis
by: Canessa, E., et al.
Published: (2023)
by: Canessa, E., et al.
Published: (2023)
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
by: Xu, Zhen, et al.
Published: (2024)
by: Xu, Zhen, et al.
Published: (2024)
Break-for-Make: Modular Low-Rank Adaptations for Composable Content-Style Customization
by: Xu, Yu, et al.
Published: (2024)
by: Xu, Yu, et al.
Published: (2024)
Similar Items
-
VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction
by: Rakesh, Vineet Kumar, et al.
Published: (2026) -
BayesFusion-SDF: Probabilistic Signed Distance Fusion with View Planning on CPU
by: Mazumdar, Soumya, et al.
Published: (2026) -
VedicTHG: Symbolic Vedic Computation for Low-Resource Talking-Head Generation in Educational Avatars
by: Rakesh, Vineet Kumar, et al.
Published: (2026) -
PrivFedTalk: Privacy-Aware Federated Diffusion with Identity-Stable Adapters for Personalized Talking-Head Generation
by: Mazumdar, Soumya, et al.
Published: (2026) -
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
by: Flynn, John, et al.
Published: (2026)