LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hao, Bowen, Zhou, Dongliang, Li, Xiaojie, Zhang, Xingyu, Xie, Liang, Wu, Jianlong, Yin, Erwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
von: Wu, Linzhi, et al.
Veröffentlicht: (2024)
von: Wu, Linzhi, et al.
Veröffentlicht: (2024)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
von: Wang, Sen, et al.
Veröffentlicht: (2025)
von: Wang, Sen, et al.
Veröffentlicht: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
Towards Accurate Lip-to-Speech Synthesis in-the-Wild
von: Hegde, Sindhu, et al.
Veröffentlicht: (2024)
von: Hegde, Sindhu, et al.
Veröffentlicht: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
Style-Preserving Lip Sync via Audio-Aware Style Reference
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
VCoME: Verbal Video Composition with Multimodal Editing Effects
von: Gong, Weibo, et al.
Veröffentlicht: (2024)
von: Gong, Weibo, et al.
Veröffentlicht: (2024)
High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
von: Ling, Zeyu, et al.
Veröffentlicht: (2025)
von: Ling, Zeyu, et al.
Veröffentlicht: (2025)
AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
von: Yin, Yongkang, et al.
Veröffentlicht: (2023)
von: Yin, Yongkang, et al.
Veröffentlicht: (2023)
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
von: Gu, Ke, et al.
Veröffentlicht: (2025)
von: Gu, Ke, et al.
Veröffentlicht: (2025)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2023)
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2023)
BC-GAN: A Generative Adversarial Network for Synthesizing a Batch of Collocated Clothing
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
von: Zhao, Jinghua, et al.
Veröffentlicht: (2025)
von: Zhao, Jinghua, et al.
Veröffentlicht: (2025)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
von: Li, Huilai, et al.
Veröffentlicht: (2026)
von: Li, Huilai, et al.
Veröffentlicht: (2026)
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
von: Li, Minghan, et al.
Veröffentlicht: (2026)
von: Li, Minghan, et al.
Veröffentlicht: (2026)
Face Consistency Benchmark for GenAI Video
von: Podstawski, Michal, et al.
Veröffentlicht: (2025)
von: Podstawski, Michal, et al.
Veröffentlicht: (2025)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
Enhancing Generalization in Medical Visual Question Answering Tasks via Gradient-Guided Model Perturbation
von: Liu, Gang, et al.
Veröffentlicht: (2024)
von: Liu, Gang, et al.
Veröffentlicht: (2024)
FCBoost-Net: A Generative Network for Synthesizing Multiple Collocated Outfits via Fashion Compatibility Boosting
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Extending Visual Dynamics for Video-to-Music Generation
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024)
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
SNP-S3: Shared Network Pre-training and Significant Semantic Strengthening for Various Video-Text Tasks
von: Dong, Xingning, et al.
Veröffentlicht: (2024)
von: Dong, Xingning, et al.
Veröffentlicht: (2024)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates
von: Wang, Zhitao, et al.
Veröffentlicht: (2025)
von: Wang, Zhitao, et al.
Veröffentlicht: (2025)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework
von: Zhang, Suoxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Suoxiang, et al.
Veröffentlicht: (2025)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
von: Li, Huilai, et al.
Veröffentlicht: (2025)
von: Li, Huilai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
von: Wu, Linzhi, et al.
Veröffentlicht: (2024) -
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
von: Wang, Sen, et al.
Veröffentlicht: (2025) -
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025) -
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024) -
Towards Accurate Lip-to-Speech Synthesis in-the-Wild
von: Hegde, Sindhu, et al.
Veröffentlicht: (2024)