Style-Preserving Lip Sync via Audio-Aware Style Reference
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhong, Weizhi, Li, Jichang, Cai, Yinqi, Li, Ming, Gao, Feng, Lin, Liang, Li, Guanbin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model
por: Zhong, Weizhi, et al.
Publicado: (2024)
por: Zhong, Weizhi, et al.
Publicado: (2024)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
por: Guan, Jiazhi, et al.
Publicado: (2024)
por: Guan, Jiazhi, et al.
Publicado: (2024)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
por: Wu, Yi, et al.
Publicado: (2025)
por: Wu, Yi, et al.
Publicado: (2025)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
por: Hsiao, Teng-Fang, et al.
Publicado: (2024)
por: Hsiao, Teng-Fang, et al.
Publicado: (2024)
StyleLipSync: Style-based Personalized Lip-sync Video Generation
por: Ki, Taekyung, et al.
Publicado: (2023)
por: Ki, Taekyung, et al.
Publicado: (2023)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
por: Li, Guangyao, et al.
Publicado: (2024)
por: Li, Guangyao, et al.
Publicado: (2024)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
por: Hao, Bowen, et al.
Publicado: (2025)
por: Hao, Bowen, et al.
Publicado: (2025)
DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
por: Hu, Ying, et al.
Publicado: (2024)
por: Hu, Ying, et al.
Publicado: (2024)
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
por: Wu, Linzhi, et al.
Publicado: (2024)
por: Wu, Linzhi, et al.
Publicado: (2024)
TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
por: Jin, Xin, et al.
Publicado: (2026)
por: Jin, Xin, et al.
Publicado: (2026)
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
por: Wang, Xuanchen, et al.
Publicado: (2025)
por: Wang, Xuanchen, et al.
Publicado: (2025)
Decoupled Audio-Visual Dataset Distillation
por: Li, Wenyuan, et al.
Publicado: (2025)
por: Li, Wenyuan, et al.
Publicado: (2025)
CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
por: Le, Anh-Duy, et al.
Publicado: (2026)
por: Le, Anh-Duy, et al.
Publicado: (2026)
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
por: Cong, Gaoxiang, et al.
Publicado: (2026)
por: Cong, Gaoxiang, et al.
Publicado: (2026)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
por: Cai, Dongnuan, et al.
Publicado: (2026)
por: Cai, Dongnuan, et al.
Publicado: (2026)
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
por: Li, Yinqi, et al.
Publicado: (2025)
por: Li, Yinqi, et al.
Publicado: (2025)
FakeRadar: Probing Forgery Outliers to Detect Unknown Deepfake Videos
por: Li, Zhaolun, et al.
Publicado: (2025)
por: Li, Zhaolun, et al.
Publicado: (2025)
COutfitGAN: Learning to Synthesize Compatible Outfits Supervised by Silhouette Masks and Fashion Styles
por: Zhou, Dongliang, et al.
Publicado: (2025)
por: Zhou, Dongliang, et al.
Publicado: (2025)
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
por: Huang, Feizhen, et al.
Publicado: (2025)
por: Huang, Feizhen, et al.
Publicado: (2025)
Break-for-Make: Modular Low-Rank Adaptations for Composable Content-Style Customization
por: Xu, Yu, et al.
Publicado: (2024)
por: Xu, Yu, et al.
Publicado: (2024)
Harnessing the Latent Diffusion Model for Training-Free Image Style Transfer
por: Masui, Kento, et al.
Publicado: (2024)
por: Masui, Kento, et al.
Publicado: (2024)
Protégé: Learn and Generate Basic Makeup Styles with Generative Adversarial Networks (GANs)
por: Sii, Jia Wei, et al.
Publicado: (2024)
por: Sii, Jia Wei, et al.
Publicado: (2024)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
por: Zhou, Jinxing, et al.
Publicado: (2024)
por: Zhou, Jinxing, et al.
Publicado: (2024)
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
por: Liao, Junchao, et al.
Publicado: (2026)
por: Liao, Junchao, et al.
Publicado: (2026)
SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment
por: Mao, Xinyu, et al.
Publicado: (2025)
por: Mao, Xinyu, et al.
Publicado: (2025)
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
por: Li, Bingzhou, et al.
Publicado: (2026)
por: Li, Bingzhou, et al.
Publicado: (2026)
AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
por: Shahzad, Sahibzada Adil, et al.
Publicado: (2023)
por: Shahzad, Sahibzada Adil, et al.
Publicado: (2023)
ROI-Guided Point Cloud Geometry Compression Towards Human and Machine Vision
por: Liang, Xie, et al.
Publicado: (2025)
por: Liang, Xie, et al.
Publicado: (2025)
Reference-Guided Identity Preserving Face Restoration
por: Zhou, Mo, et al.
Publicado: (2025)
por: Zhou, Mo, et al.
Publicado: (2025)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
por: Zhang, Jiaxu, et al.
Publicado: (2025)
por: Zhang, Jiaxu, et al.
Publicado: (2025)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
por: Qi, Xingqun, et al.
Publicado: (2023)
por: Qi, Xingqun, et al.
Publicado: (2023)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
por: Ling, Jun, et al.
Publicado: (2024)
por: Ling, Jun, et al.
Publicado: (2024)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
por: Li, Huilai, et al.
Publicado: (2025)
por: Li, Huilai, et al.
Publicado: (2025)
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
por: Zhang, Zhicheng, et al.
Publicado: (2026)
por: Zhang, Zhicheng, et al.
Publicado: (2026)
PAND: Prompt-Aware Neighborhood Distillation for Lightweight Fine-Grained Visual Classification
por: Luo, Qiuming, et al.
Publicado: (2026)
por: Luo, Qiuming, et al.
Publicado: (2026)
StereoSync: Spatially-Aware Stereo Audio Generation from Video
por: Marinoni, Christian, et al.
Publicado: (2025)
por: Marinoni, Christian, et al.
Publicado: (2025)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
por: Tian, Zeyue, et al.
Publicado: (2026)
por: Tian, Zeyue, et al.
Publicado: (2026)
ToonAging: Face Re-Aging upon Artistic Portrait Style Transfer
por: Kim, Bumsoo, et al.
Publicado: (2024)
por: Kim, Bumsoo, et al.
Publicado: (2024)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
por: Huang, Victor Shea-Jay, et al.
Publicado: (2025)
por: Huang, Victor Shea-Jay, et al.
Publicado: (2025)
SonoWorld: From One Image to a 3D Audio-Visual Scene
por: Jin, Derong, et al.
Publicado: (2026)
por: Jin, Derong, et al.
Publicado: (2026)
Ejemplares similares
-
High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model
por: Zhong, Weizhi, et al.
Publicado: (2024) -
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
por: Guan, Jiazhi, et al.
Publicado: (2024) -
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
por: Wu, Yi, et al.
Publicado: (2025) -
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
por: Hsiao, Teng-Fang, et al.
Publicado: (2024) -
StyleLipSync: Style-based Personalized Lip-sync Video Generation
por: Ki, Taekyung, et al.
Publicado: (2023)