M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Kui, Liu, Shiyu, Jiang, Junjun, Yao, Hongxun, Fan, Xiaopeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis
von: Liu, Shiyu, et al.
Veröffentlicht: (2025)
von: Liu, Shiyu, et al.
Veröffentlicht: (2025)
Dynamic Policy-Driven Adaptive Multi-Instance Learning for Whole Slide Image Classification
von: Zheng, Tingting, et al.
Veröffentlicht: (2024)
von: Zheng, Tingting, et al.
Veröffentlicht: (2024)
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
von: Liu, Tao, et al.
Veröffentlicht: (2024)
von: Liu, Tao, et al.
Veröffentlicht: (2024)
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
von: Zhong, Zhizhou, et al.
Veröffentlicht: (2025)
von: Zhong, Zhizhou, et al.
Veröffentlicht: (2025)
DAM: Dual Active Learning with Multimodal Foundation Model for Source-Free Domain Adaptation
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation
von: Kim, Seyeon, et al.
Veröffentlicht: (2024)
von: Kim, Seyeon, et al.
Veröffentlicht: (2024)
FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust Fusion
von: Sun, Pihai, et al.
Veröffentlicht: (2025)
von: Sun, Pihai, et al.
Veröffentlicht: (2025)
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
von: Wang, Mengchao, et al.
Veröffentlicht: (2025)
von: Wang, Mengchao, et al.
Veröffentlicht: (2025)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
von: Shen, Yijun, et al.
Veröffentlicht: (2025)
von: Shen, Yijun, et al.
Veröffentlicht: (2025)
SfMamba: Efficient Source-Free Domain Adaptation via Selective Scan Modeling
von: Chen, Xi, et al.
Veröffentlicht: (2026)
von: Chen, Xi, et al.
Veröffentlicht: (2026)
Learning from History: Task-agnostic Model Contrastive Learning for Image Restoration
von: Wu, Gang, et al.
Veröffentlicht: (2023)
von: Wu, Gang, et al.
Veröffentlicht: (2023)
Boosting All-in-One Image Restoration via Self-Improved Privilege Learning
von: Wu, Gang, et al.
Veröffentlicht: (2025)
von: Wu, Gang, et al.
Veröffentlicht: (2025)
EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans
von: Zhou, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhou, Yingjie, et al.
Veröffentlicht: (2025)
Improving Domain Generalization in Self-supervised Monocular Depth Estimation via Stabilized Adversarial Training
von: Yao, Yuanqi, et al.
Veröffentlicht: (2024)
von: Yao, Yuanqi, et al.
Veröffentlicht: (2024)
D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head Synthesis
von: Guo, Yuhang, et al.
Veröffentlicht: (2025)
von: Guo, Yuhang, et al.
Veröffentlicht: (2025)
Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
von: Zhou, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhou, Yingjie, et al.
Veröffentlicht: (2025)
Fast and Accurate Gigapixel Pathological Image Classification with Hierarchical Distillation Multi-Instance Learning
von: Dong, Jiuyang, et al.
Veröffentlicht: (2025)
von: Dong, Jiuyang, et al.
Veröffentlicht: (2025)
AHDMIL: Asymmetric Hierarchical Distillation Multi-Instance Learning for Fast and Accurate Whole-Slide Image Classification
von: Dong, Jiuyang, et al.
Veröffentlicht: (2025)
von: Dong, Jiuyang, et al.
Veröffentlicht: (2025)
Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
von: Li, Wenrui, et al.
Veröffentlicht: (2025)
von: Li, Wenrui, et al.
Veröffentlicht: (2025)
Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image Deraining
von: Yu, Zhaocheng, et al.
Veröffentlicht: (2025)
von: Yu, Zhaocheng, et al.
Veröffentlicht: (2025)
LLV-FSR: Exploiting Large Language-Vision Prior for Face Super-resolution
von: Wang, Chenyang, et al.
Veröffentlicht: (2024)
von: Wang, Chenyang, et al.
Veröffentlicht: (2024)
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
von: Yao, Ziyu, et al.
Veröffentlicht: (2024)
von: Yao, Ziyu, et al.
Veröffentlicht: (2024)
Fully $1\times1$ Convolutional Network for Lightweight Image Super-Resolution
von: Wu, Gang, et al.
Veröffentlicht: (2023)
von: Wu, Gang, et al.
Veröffentlicht: (2023)
Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art Style
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space
von: Feng, Guanwen, et al.
Veröffentlicht: (2024)
von: Feng, Guanwen, et al.
Veröffentlicht: (2024)
DashGaussian: Optimizing 3D Gaussian Splatting in 200 Seconds
von: Chen, Youyu, et al.
Veröffentlicht: (2025)
von: Chen, Youyu, et al.
Veröffentlicht: (2025)
DSwinIR: Rethinking Window-based Attention for Image Restoration
von: Wu, Gang, et al.
Veröffentlicht: (2025)
von: Wu, Gang, et al.
Veröffentlicht: (2025)
COB-GS: Clear Object Boundaries in 3DGS Segmentation Based on Boundary-Adaptive Gaussian Splitting
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2025)
Beyond Degradation Redundancy: Contrastive Prompt Learning for All-in-One Image Restoration
von: Wu, Gang, et al.
Veröffentlicht: (2025)
von: Wu, Gang, et al.
Veröffentlicht: (2025)
SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
von: Xiong, Lingyu, et al.
Veröffentlicht: (2024)
von: Xiong, Lingyu, et al.
Veröffentlicht: (2024)
GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
FrozenSeg: Harmonizing Frozen Foundation Models for Open-Vocabulary Segmentation
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
von: Feng, Guanwen, et al.
Veröffentlicht: (2025)
von: Feng, Guanwen, et al.
Veröffentlicht: (2025)
Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios
von: Xu, Jialei, et al.
Veröffentlicht: (2024)
von: Xu, Jialei, et al.
Veröffentlicht: (2024)
Beyond Talking -- Generating Holistic 3D Human Dyadic Motion for Communication
von: Sun, Mingze, et al.
Veröffentlicht: (2024)
von: Sun, Mingze, et al.
Veröffentlicht: (2024)
Exploiting Self-Supervised Constraints in Image Super-Resolution
von: Wu, Gang, et al.
Veröffentlicht: (2024)
von: Wu, Gang, et al.
Veröffentlicht: (2024)
OpFlowTalker: Realistic and Natural Talking Face Generation via Optical Flow Guidance
von: Ge, Shuheng, et al.
Veröffentlicht: (2024)
von: Ge, Shuheng, et al.
Veröffentlicht: (2024)
Reliev3R: Relieving Feed-forward Reconstruction from Multi-View Geometric Annotations
von: Chen, Youyu, et al.
Veröffentlicht: (2026)
von: Chen, Youyu, et al.
Veröffentlicht: (2026)
Spatial Annealing for Efficient Few-shot Neural Rendering
von: Xiao, Yuru, et al.
Veröffentlicht: (2024)
von: Xiao, Yuru, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis
von: Liu, Shiyu, et al.
Veröffentlicht: (2025) -
Dynamic Policy-Driven Adaptive Multi-Instance Learning for Whole Slide Image Classification
von: Zheng, Tingting, et al.
Veröffentlicht: (2024) -
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
von: Liu, Tao, et al.
Veröffentlicht: (2024) -
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
von: Zhong, Zhizhou, et al.
Veröffentlicht: (2025) -
DAM: Dual Active Learning with Multimodal Foundation Model for Source-Free Domain Adaptation
von: Chen, Xi, et al.
Veröffentlicht: (2025)