PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ling, Jun, Wang, Yiwen, Xue, Han, Xie, Rong, Song, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
von: Tang, Anni, et al.
Veröffentlicht: (2022)
von: Tang, Anni, et al.
Veröffentlicht: (2022)
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026)
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026)
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
von: Yao, Ziyu, et al.
Veröffentlicht: (2024)
von: Yao, Ziyu, et al.
Veröffentlicht: (2024)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
von: Chen, Sen, et al.
Veröffentlicht: (2022)
von: Chen, Sen, et al.
Veröffentlicht: (2022)
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
von: Zhang, Beiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Beiyuan, et al.
Veröffentlicht: (2024)
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
von: Flynn, John, et al.
Veröffentlicht: (2026)
von: Flynn, John, et al.
Veröffentlicht: (2026)
GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting
von: Cho, Kyusun, et al.
Veröffentlicht: (2024)
von: Cho, Kyusun, et al.
Veröffentlicht: (2024)
GAIA: Zero-shot Talking Avatar Generation
von: He, Tianyu, et al.
Veröffentlicht: (2023)
von: He, Tianyu, et al.
Veröffentlicht: (2023)
TIPS: Text-Induced Pose Synthesis
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise Control
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
von: Xiong, Lingyu, et al.
Veröffentlicht: (2024)
von: Xiong, Lingyu, et al.
Veröffentlicht: (2024)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression Editing
von: Jiang, Diqiong, et al.
Veröffentlicht: (2026)
von: Jiang, Diqiong, et al.
Veröffentlicht: (2026)
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait
von: Ki, Taekyung, et al.
Veröffentlicht: (2024)
von: Ki, Taekyung, et al.
Veröffentlicht: (2024)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial Animation
von: Chen, Liyang, et al.
Veröffentlicht: (2023)
von: Chen, Liyang, et al.
Veröffentlicht: (2023)
OpFlowTalker: Realistic and Natural Talking Face Generation via Optical Flow Guidance
von: Ge, Shuheng, et al.
Veröffentlicht: (2024)
von: Ge, Shuheng, et al.
Veröffentlicht: (2024)
NeRF-AD: Neural Radiance Field with Attention-based Disentanglement for Talking Face Synthesis
von: Bi, Chongke, et al.
Veröffentlicht: (2024)
von: Bi, Chongke, et al.
Veröffentlicht: (2024)
G-Refine: A General Quality Refiner for Text-to-Image Generation
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
A Near-Raw Talking-Head Video Dataset for Various Computer Vision Tasks
von: Naderi, Babak, et al.
Veröffentlicht: (2026)
von: Naderi, Babak, et al.
Veröffentlicht: (2026)
Multi-scale Attention Guided Pose Transfer
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
von: Liu, Ke, et al.
Veröffentlicht: (2026)
von: Liu, Ke, et al.
Veröffentlicht: (2026)
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
von: Li, Dong, et al.
Veröffentlicht: (2025)
von: Li, Dong, et al.
Veröffentlicht: (2025)
VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026)
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026)
DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
von: Dong, Ziyi, et al.
Veröffentlicht: (2022)
von: Dong, Ziyi, et al.
Veröffentlicht: (2022)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
von: Taghipour, Ashkan, et al.
Veröffentlicht: (2026)
von: Taghipour, Ashkan, et al.
Veröffentlicht: (2026)
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
Is It Really You? Exploring Biometric Verification Scenarios in Photorealistic Talking-Head Avatar Videos
von: Pedrouzo-Rodriguez, Laura, et al.
Veröffentlicht: (2025)
von: Pedrouzo-Rodriguez, Laura, et al.
Veröffentlicht: (2025)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
Detection and Recovery of Adversarial Slow-Pose Drift in Offloaded Visual-Inertial Odometry
von: Saha, Soruya, et al.
Veröffentlicht: (2025)
von: Saha, Soruya, et al.
Veröffentlicht: (2025)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
von: Kang, Fang, et al.
Veröffentlicht: (2025)
von: Kang, Fang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
von: Tang, Anni, et al.
Veröffentlicht: (2022) -
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026) -
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
von: Yao, Ziyu, et al.
Veröffentlicht: (2024) -
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026) -
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
von: Chen, Sen, et al.
Veröffentlicht: (2022)