FluentLip: A Phonemes-Based Two-stage Approach for Audio-Driven Lip Synthesis with Optical Flow Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shiyan, Qu, Rui, Jin, Yan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeepSORT-Driven Visual Tracking Approach for Gesture Recognition in Interactive Systems
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
Automated Image-Based Identification and Consistent Classification of Fire Patterns with Quantitative Shape Analysis and Spatial Location Identification
by: Liu, Pengkun, et al.
Published: (2024)
by: Liu, Pengkun, et al.
Published: (2024)
Visual Neural Decoding via Improved Visual-EEG Semantic Consistency
by: Chen, Hongzhou, et al.
Published: (2024)
by: Chen, Hongzhou, et al.
Published: (2024)
Towards Context-aware Support for Color Vision Deficiency: An Approach Integrating LLM and AR
by: Morita, Shogo, et al.
Published: (2024)
by: Morita, Shogo, et al.
Published: (2024)
Low Latency Gaze Tracking via Latent Optical Sensing
by: Zheng, Yidan, et al.
Published: (2026)
by: Zheng, Yidan, et al.
Published: (2026)
An Efficient and Streaming Audio Visual Active Speaker Detection System
by: Kundu, Arnav, et al.
Published: (2024)
by: Kundu, Arnav, et al.
Published: (2024)
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
by: Park, Se Jin, et al.
Published: (2024)
by: Park, Se Jin, et al.
Published: (2024)
Generalized Pose Space Embeddings for Training In-the-Wild using Anaylis-by-Synthesis
by: Borer, Dominik, et al.
Published: (2024)
by: Borer, Dominik, et al.
Published: (2024)
Efficient Listener: Dyadic Facial Motion Synthesis via Action Diffusion
by: Wang, Zesheng, et al.
Published: (2025)
by: Wang, Zesheng, et al.
Published: (2025)
SketchPlay: Intuitive Creation of Physically Realistic VR Content with Gesture-Driven Sketching
by: Zhang, Xiangwen, et al.
Published: (2025)
by: Zhang, Xiangwen, et al.
Published: (2025)
VizDefender: Unmasking Visualization Tampering through Proactive Localization and Intent Inference
by: Song, Sicheng, et al.
Published: (2025)
by: Song, Sicheng, et al.
Published: (2025)
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
by: Xu, Zunnan, et al.
Published: (2024)
by: Xu, Zunnan, et al.
Published: (2024)
Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention
by: Chang, Minsuk, et al.
Published: (2025)
by: Chang, Minsuk, et al.
Published: (2025)
WheelPose: Data Synthesis Techniques to Improve Pose Estimation Performance on Wheelchair Users
by: Huang, William, et al.
Published: (2024)
by: Huang, William, et al.
Published: (2024)
DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis
by: Deng, Kaijun, et al.
Published: (2024)
by: Deng, Kaijun, et al.
Published: (2024)
Uncovering the Metaverse within Everyday Environments: a Coarse-to-Fine Approach
by: Xu, Liming, et al.
Published: (2024)
by: Xu, Liming, et al.
Published: (2024)
PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis
by: Huang, Yihuan, et al.
Published: (2025)
by: Huang, Yihuan, et al.
Published: (2025)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
by: Liu, Xinyi, et al.
Published: (2025)
by: Liu, Xinyi, et al.
Published: (2025)
Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakes
by: Liu, Weifeng, et al.
Published: (2024)
by: Liu, Weifeng, et al.
Published: (2024)
Human Motion Synthesis_ A Diffusion Approach for Motion Stitching and In-Betweening
by: Adewole, Michael, et al.
Published: (2024)
by: Adewole, Michael, et al.
Published: (2024)
PoseDriver: A Unified Approach to Multi-Category Skeleton Detection for Autonomous Driving
by: Borhani, Yasamin, et al.
Published: (2026)
by: Borhani, Yasamin, et al.
Published: (2026)
Fuzzy Logic Approach For Visual Analysis Of Websites With K-means Clustering-based Color Extraction
by: Abildayeva, Tamiris, et al.
Published: (2024)
by: Abildayeva, Tamiris, et al.
Published: (2024)
Blind Augmentation: Calibration-free Camera Distortion Model Estimation for Real-time Mixed-reality Consistency
by: Prakash, Siddhant, et al.
Published: (2025)
by: Prakash, Siddhant, et al.
Published: (2025)
ScrollTimes: Tracing the Provenance of Paintings as a Window into History
by: Zhang, Wei, et al.
Published: (2023)
by: Zhang, Wei, et al.
Published: (2023)
Evaluating Sensitivity Parameters in Smartphone-Based Gaze Estimation: A Comparative Study of Appearance-Based and Infrared Eye Trackers
by: Gunawardena, Nishan, et al.
Published: (2025)
by: Gunawardena, Nishan, et al.
Published: (2025)
LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping Deformation
by: Junli, Deng, et al.
Published: (2024)
by: Junli, Deng, et al.
Published: (2024)
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
by: Ma, Junxian, et al.
Published: (2025)
by: Ma, Junxian, et al.
Published: (2025)
Leveraging Digital Perceptual Technologies for Remote Perception and Analysis of Human Biomechanical Processes: A Contactless Approach for Workload and Joint Force Assessment
by: Omidokun, Jesudara, et al.
Published: (2024)
by: Omidokun, Jesudara, et al.
Published: (2024)
Impact of Target and Tool Visualization on Depth Perception and Usability in Optical See-Through AR
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
A Comparative Study of Scanpath Models in Graph-Based Visualization
by: Lopez-Cardona, Angela, et al.
Published: (2025)
by: Lopez-Cardona, Angela, et al.
Published: (2025)
Machine Vision-Based Surgical Lighting System:Design and Implementation
by: Gharghabi, Amir, et al.
Published: (2025)
by: Gharghabi, Amir, et al.
Published: (2025)
Breaking Coordinate Overfitting: Geometry-Aware WiFi Sensing for Cross-Layout 3D Pose Estimation
by: Jia, Songming, et al.
Published: (2026)
by: Jia, Songming, et al.
Published: (2026)
A CNN Based Framework for Unistroke Numeral Recognition in Air-Writing
by: Roy, Prasun, et al.
Published: (2023)
by: Roy, Prasun, et al.
Published: (2023)
GroundUp: Rapid Sketch-Based 3D City Massing
by: Unlu, Gizem Esra, et al.
Published: (2024)
by: Unlu, Gizem Esra, et al.
Published: (2024)
iTrace: Click-Based Gaze Visualization on the Apple Vision Pro
by: Mehmedova, Esra, et al.
Published: (2025)
by: Mehmedova, Esra, et al.
Published: (2025)
MIBURI: Towards Expressive Interactive Gesture Synthesis
by: Mughal, M. Hamza, et al.
Published: (2026)
by: Mughal, M. Hamza, et al.
Published: (2026)
Computational Trichromacy Reconstruction: Empowering the Color-Vision Deficient to Recognize Colors Using Augmented Reality
by: Zhu, Yuhao, et al.
Published: (2024)
by: Zhu, Yuhao, et al.
Published: (2024)
Night Eyes: A Reproducible Framework for Constellation-Based Corneal Reflection Matching
by: Maquiling, Virmarie, et al.
Published: (2026)
by: Maquiling, Virmarie, et al.
Published: (2026)
Machine Learning-Based Jamun Leaf Disease Detection: A Comprehensive Review
by: Bhowmik, Auvick Chandra, et al.
Published: (2023)
by: Bhowmik, Auvick Chandra, et al.
Published: (2023)
A Convolution-Based Gait Asymmetry Metric for Inter-Limb Synergistic Coordination
by: Fukino, Go, et al.
Published: (2025)
by: Fukino, Go, et al.
Published: (2025)
Similar Items
-
DeepSORT-Driven Visual Tracking Approach for Gesture Recognition in Interactive Systems
by: Zhang, Tong, et al.
Published: (2025) -
Automated Image-Based Identification and Consistent Classification of Fire Patterns with Quantitative Shape Analysis and Spatial Location Identification
by: Liu, Pengkun, et al.
Published: (2024) -
Visual Neural Decoding via Improved Visual-EEG Semantic Consistency
by: Chen, Hongzhou, et al.
Published: (2024) -
Towards Context-aware Support for Color Vision Deficiency: An Approach Integrating LLM and AR
by: Morita, Shogo, et al.
Published: (2024) -
Low Latency Gaze Tracking via Latent Optical Sensing
by: Zheng, Yidan, et al.
Published: (2026)