VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Symeonidis-Herzig, Alexandre, Sincan, Özge Mercanoğlu, Bowden, Richard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production
by: Symeonidis-Herzig, Alexandre, et al.
Published: (2026)
by: Symeonidis-Herzig, Alexandre, et al.
Published: (2026)
Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation
by: Sincan, Ozge Mercanoglu, et al.
Published: (2025)
by: Sincan, Ozge Mercanoglu, et al.
Published: (2025)
Sign Spotting Disambiguation using Large Language Models
by: Low, JianHe, et al.
Published: (2025)
by: Low, JianHe, et al.
Published: (2025)
SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning
by: Low, Jianhe, et al.
Published: (2026)
by: Low, Jianhe, et al.
Published: (2026)
Spotter+GPT: Turning Sign Spottings into Sentences with LLMs
by: Sincan, Ozge Mercanoglu, et al.
Published: (2024)
by: Sincan, Ozge Mercanoglu, et al.
Published: (2024)
Hands-On: Segmenting Individual Signs from Continuous Sequences
by: Low, JianHe, et al.
Published: (2025)
by: Low, JianHe, et al.
Published: (2025)
SignAgent: Agentic LLMs for Linguistically-Grounded Sign Language Annotation and Dataset Curation
by: Cory, Oliver, et al.
Published: (2026)
by: Cory, Oliver, et al.
Published: (2026)
Giving a Hand to Diffusion Models: a Two-Stage Approach to Improving Conditional Human Image Generation
by: Pelykh, Anton, et al.
Published: (2024)
by: Pelykh, Anton, et al.
Published: (2024)
SAGE: Segment-Aware Gloss-Free Encoding for Token-Efficient Sign Language Translation
by: Low, JianHe, et al.
Published: (2025)
by: Low, JianHe, et al.
Published: (2025)
Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation
by: Asasi, Sobhan, et al.
Published: (2025)
by: Asasi, Sobhan, et al.
Published: (2025)
Gloss-Free Sign Language Translation: An Unbiased Evaluation of Progress in the Field
by: Sincan, Ozge Mercanoglu, et al.
Published: (2026)
by: Sincan, Ozge Mercanoglu, et al.
Published: (2026)
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
by: Wu, Linzhi, et al.
Published: (2024)
by: Wu, Linzhi, et al.
Published: (2024)
VALLR: Visual ASR Language Model for Lip Reading
by: Thomas, Marshall, et al.
Published: (2025)
by: Thomas, Marshall, et al.
Published: (2025)
NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition
by: Yao, Junguang, et al.
Published: (2026)
by: Yao, Junguang, et al.
Published: (2026)
Modelling the Distribution of Human Motion for Sign Language Assessment
by: Cory, Oliver, et al.
Published: (2024)
by: Cory, Oliver, et al.
Published: (2024)
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
AvatarShield: Visual Reinforcement Learning for Human-Centric Synthetic Video Detection
by: Xu, Zhipei, et al.
Published: (2025)
by: Xu, Zhipei, et al.
Published: (2025)
STNet: Deep Audio-Visual Fusion Network for Robust Speaker Tracking
by: Li, Yidi, et al.
Published: (2024)
by: Li, Yidi, et al.
Published: (2024)
InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
by: Wang, Yuchi, et al.
Published: (2024)
by: Wang, Yuchi, et al.
Published: (2024)
PuzzleAvatar: Assembling 3D Avatars from Personal Albums
by: Xiu, Yuliang, et al.
Published: (2024)
by: Xiu, Yuliang, et al.
Published: (2024)
Reasoning Matters for 3D Visual Grounding
by: Huang, Hsiang-Wei, et al.
Published: (2026)
by: Huang, Hsiang-Wei, et al.
Published: (2026)
Text2Avatar: Text to 3D Human Avatar Generation with Codebook-Driven Body Controllable Attribute
by: Gong, Chaoqun, et al.
Published: (2024)
by: Gong, Chaoqun, et al.
Published: (2024)
Learning Separable Hidden Unit Contributions for Speaker-Adaptive Lip-Reading
by: Luo, Songtao, et al.
Published: (2023)
by: Luo, Songtao, et al.
Published: (2023)
AEGIS: Preserving privacy of 3D Facial Avatars with Adversarial Perturbations
by: Wolkiewicz, Dawid, et al.
Published: (2025)
by: Wolkiewicz, Dawid, et al.
Published: (2025)
Saliency Guided Longitudinal Medical Visual Question Answering
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
F3G-Avatar : Face Focused Full-body Gaussian Avatar
by: Menu, Willem, et al.
Published: (2026)
by: Menu, Willem, et al.
Published: (2026)
Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-Attention
by: Leem, Saebom, et al.
Published: (2024)
by: Leem, Saebom, et al.
Published: (2024)
AMG: Avatar Motion Guided Video Generation
by: Yang, Zhangsihao, et al.
Published: (2024)
by: Yang, Zhangsihao, et al.
Published: (2024)
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
by: Lin, Jiawen, et al.
Published: (2025)
by: Lin, Jiawen, et al.
Published: (2025)
FluentLip: A Phonemes-Based Two-stage Approach for Audio-Driven Lip Synthesis with Optical Flow Consistency
by: Liu, Shiyan, et al.
Published: (2025)
by: Liu, Shiyan, et al.
Published: (2025)
TexAvatars : Hybrid Texel-3D Representations for Stable Rigging of Photorealistic Gaussian Head Avatars
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
GPTDrawer: Enhancing Visual Synthesis through ChatGPT
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Towards High-fidelity 3D Talking Avatar with Personalized Dynamic Texture
by: Li, Xuanchen, et al.
Published: (2025)
by: Li, Xuanchen, et al.
Published: (2025)
Morphable Diffusion: 3D-Consistent Diffusion for Single-image Avatar Creation
by: Chen, Xiyi, et al.
Published: (2024)
by: Chen, Xiyi, et al.
Published: (2024)
Audio-Guided Visual Perception for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
L3D-Pose: Lifting Pose for 3D Avatars from a Single Camera in the Wild
by: Debnath, Soumyaratna, et al.
Published: (2025)
by: Debnath, Soumyaratna, et al.
Published: (2025)
Tamaththul3D: High-Fidelity 3D Saudi Sign Language Avatars from Monocular Video
by: Alghamdi, Eyad, et al.
Published: (2026)
by: Alghamdi, Eyad, et al.
Published: (2026)
Mon3tr: Monocular 3D Telepresence with Pre-built Gaussian Avatars as Amortization
by: Lin, Fangyu, et al.
Published: (2026)
by: Lin, Fangyu, et al.
Published: (2026)
VERSE: Visual Embedding Reduction and Space Exploration. Clustering-Guided Insights for Training Data Enhancement in Visually-Rich Document Understanding
by: de Rodrigo, Ignacio, et al.
Published: (2026)
by: de Rodrigo, Ignacio, et al.
Published: (2026)
All in One: Visual-Description-Guided Unified Point Cloud Segmentation
by: Han, Zongyan, et al.
Published: (2025)
by: Han, Zongyan, et al.
Published: (2025)
Similar Items
-
M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production
by: Symeonidis-Herzig, Alexandre, et al.
Published: (2026) -
Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation
by: Sincan, Ozge Mercanoglu, et al.
Published: (2025) -
Sign Spotting Disambiguation using Large Language Models
by: Low, JianHe, et al.
Published: (2025) -
SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning
by: Low, Jianhe, et al.
Published: (2026) -
Spotter+GPT: Turning Sign Spottings into Sentences with LLMs
by: Sincan, Ozge Mercanoglu, et al.
Published: (2024)