GAIA: Zero-shot Talking Avatar Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Tianyu, Guo, Junliang, Yu, Runyi, Wang, Yuchi, Zhu, Jialiang, An, Kaikai, Li, Leyi, Tan, Xu, Wang, Chunyu, Hu, Han, Wu, HsiangTao, Zhao, Sheng, Bian, Jiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
von: Tang, Anni, et al.
Veröffentlicht: (2022)
von: Tang, Anni, et al.
Veröffentlicht: (2022)
Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning
von: Wang, Hanyao, et al.
Veröffentlicht: (2024)
von: Wang, Hanyao, et al.
Veröffentlicht: (2024)
EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026)
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026)
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
Modularized Zero-shot VQA with Pre-trained Models
von: Cao, Rui, et al.
Veröffentlicht: (2023)
von: Cao, Rui, et al.
Veröffentlicht: (2023)
Make Your Actor Talk: Generalizable and High-Fidelity Lip Sync with Motion and Appearance Disentanglement
von: Yu, Runyi, et al.
Veröffentlicht: (2024)
von: Yu, Runyi, et al.
Veröffentlicht: (2024)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
von: Zhang, Han, et al.
Veröffentlicht: (2025)
von: Zhang, Han, et al.
Veröffentlicht: (2025)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs
von: Wu, Yike, et al.
Veröffentlicht: (2026)
von: Wu, Yike, et al.
Veröffentlicht: (2026)
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
DiffBrush:Just Painting the Art by Your Hands
von: Chu, Jiaming, et al.
Veröffentlicht: (2025)
von: Chu, Jiaming, et al.
Veröffentlicht: (2025)
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)
OmniGAIA: Towards Native Omni-Modal AI Agents
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
Hallucination Localization in Video Captioning
von: Nakada, Shota, et al.
Veröffentlicht: (2025)
von: Nakada, Shota, et al.
Veröffentlicht: (2025)
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
Is It Really You? Exploring Biometric Verification Scenarios in Photorealistic Talking-Head Avatar Videos
von: Pedrouzo-Rodriguez, Laura, et al.
Veröffentlicht: (2025)
von: Pedrouzo-Rodriguez, Laura, et al.
Veröffentlicht: (2025)
Selective Vision-Language Subspace Projection for Few-shot CLIP
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
Semantic Compensation via Adversarial Removal for Robust Zero-Shot ECG Diagnosis
von: Liu, Hongjun, et al.
Veröffentlicht: (2026)
von: Liu, Hongjun, et al.
Veröffentlicht: (2026)
Personalized Playback Technology: How Short Video Services Create Excellent User Experience
von: Deng, Weihui, et al.
Veröffentlicht: (2024)
von: Deng, Weihui, et al.
Veröffentlicht: (2024)
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection
von: Liu, Yu, et al.
Veröffentlicht: (2025)
von: Liu, Yu, et al.
Veröffentlicht: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
Talking Slide Avatars: Open-Source Multimodal Communication Approach for Teaching
von: Wu, Xinxing
Veröffentlicht: (2026)
von: Wu, Xinxing
Veröffentlicht: (2026)
TSC-PCAC: Voxel Transformer and Sparse Convolution Based Point Cloud Attribute Compression for 3D Broadcasting
von: Guo, Zixi, et al.
Veröffentlicht: (2024)
von: Guo, Zixi, et al.
Veröffentlicht: (2024)
MM-InstructEval: Zero-Shot Evaluation of (Multimodal) Large Language Models on Multimodal Reasoning Tasks
von: Yang, Xiaocui, et al.
Veröffentlicht: (2024)
von: Yang, Xiaocui, et al.
Veröffentlicht: (2024)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
von: Yang, Chenglin, et al.
Veröffentlicht: (2023)
von: Yang, Chenglin, et al.
Veröffentlicht: (2023)
Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information Extraction
von: Zhou, Baohang, et al.
Veröffentlicht: (2026)
von: Zhou, Baohang, et al.
Veröffentlicht: (2026)
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
von: Zhu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Zhu, Jiaqi, et al.
Veröffentlicht: (2024)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
Zero-shot image privacy classification with Vision-Language Models
von: Baia, Alina Elena, et al.
Veröffentlicht: (2025)
von: Baia, Alina Elena, et al.
Veröffentlicht: (2025)
2DGS-Avatar: Animatable High-fidelity Clothed Avatar via 2D Gaussian Splatting
von: Yan, Qipeng, et al.
Veröffentlicht: (2025)
von: Yan, Qipeng, et al.
Veröffentlicht: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
von: Zhou, S. Z., et al.
Veröffentlicht: (2025)
von: Zhou, S. Z., et al.
Veröffentlicht: (2025)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
von: Xie, Yifan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023) -
InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
von: Wang, Yuchi, et al.
Veröffentlicht: (2024) -
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
von: Tang, Anni, et al.
Veröffentlicht: (2022) -
Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning
von: Wang, Hanyao, et al.
Veröffentlicht: (2024) -
EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot
von: Fei, Hao, et al.
Veröffentlicht: (2024)