Vclip: Face-based Speaker Generation by Face-voice Association Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Yao, Xu, Yunfei, Suo, Hongbin, Wan, Yulong, Liu, Haifeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization
by: Tao, Ruijie, et al.
Published: (2024)
by: Tao, Ruijie, et al.
Published: (2024)
Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System
by: Li, Ze, et al.
Published: (2024)
by: Li, Ze, et al.
Published: (2024)
The Database and Benchmark for the Source Speaker Tracing Challenge 2024
by: Li, Ze, et al.
Published: (2024)
by: Li, Ze, et al.
Published: (2024)
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
by: Clarke, Jason, et al.
Published: (2025)
by: Clarke, Jason, et al.
Published: (2025)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
by: Wang, Haoyu, et al.
Published: (2022)
by: Wang, Haoyu, et al.
Published: (2022)
Hear Your Face: Face-based voice conversion with F0 estimation
by: Lee, Jaejun, et al.
Published: (2024)
by: Lee, Jaejun, et al.
Published: (2024)
Task-Agnostic Structured Pruning of Speech Representation Models
by: Wang, Haoyu, et al.
Published: (2023)
by: Wang, Haoyu, et al.
Published: (2023)
Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association
by: Chen, Wuyang, et al.
Published: (2024)
by: Chen, Wuyang, et al.
Published: (2024)
Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan
by: Saeed, Muhammad Saad, et al.
Published: (2024)
by: Saeed, Muhammad Saad, et al.
Published: (2024)
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
by: Park, Kyeongman, et al.
Published: (2025)
by: Park, Kyeongman, et al.
Published: (2025)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
by: Shao, Yiwen, et al.
Published: (2024)
by: Shao, Yiwen, et al.
Published: (2024)
Learning Emotion-Invariant Speaker Representations for Speaker Verification
by: Tian, Jingguang, et al.
Published: (2025)
by: Tian, Jingguang, et al.
Published: (2025)
IDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector Mapping
by: Liu, Zeyan, et al.
Published: (2025)
by: Liu, Zeyan, et al.
Published: (2025)
A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
by: Zhou, Zhenyu, et al.
Published: (2024)
by: Zhou, Zhenyu, et al.
Published: (2024)
MASV: Speaker Verification with Global and Local Context Mamba
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
by: Pan, Zexu, et al.
Published: (2025)
by: Pan, Zexu, et al.
Published: (2025)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
by: Fu, Ruibo, et al.
Published: (2024)
by: Fu, Ruibo, et al.
Published: (2024)
Voice Conversion Augmentation for Speaker Recognition on Defective Datasets
by: Tao, Ruijie, et al.
Published: (2024)
by: Tao, Ruijie, et al.
Published: (2024)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
by: Su, Fei, et al.
Published: (2026)
by: Su, Fei, et al.
Published: (2026)
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
by: Aldeneh, Zakaria, et al.
Published: (2024)
by: Aldeneh, Zakaria, et al.
Published: (2024)
Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice
by: Rautenberg, Frederik, et al.
Published: (2025)
by: Rautenberg, Frederik, et al.
Published: (2025)
Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
by: Weng, Shilong, et al.
Published: (2025)
by: Weng, Shilong, et al.
Published: (2025)
Target Speaker Extraction with Curriculum Learning
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
Speaker Contrastive Learning for Source Speaker Tracing
by: Wang, Qing, et al.
Published: (2024)
by: Wang, Qing, et al.
Published: (2024)
Eigenvoice Synthesis based on Model Editing for Speaker Generation
by: Murata, Masato, et al.
Published: (2025)
by: Murata, Masato, et al.
Published: (2025)
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Unsupervised Face-Masked Speech Enhancement Using Generative Adversarial Networks With Human-in-the-Loop Assessment Metrics
by: Wang, Syu-Siang, et al.
Published: (2024)
by: Wang, Syu-Siang, et al.
Published: (2024)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
by: Chen, Zhengyang, et al.
Published: (2024)
by: Chen, Zhengyang, et al.
Published: (2024)
Study on Inter and Intra Speaker Variability in Speaker Recognition
by: Okhotnikov, Anton, et al.
Published: (2024)
by: Okhotnikov, Anton, et al.
Published: (2024)
Magnitude and Phase-based Feature Fusion Using Co-attention Mechanism for Speaker recognition
by: Su, Rongfeng, et al.
Published: (2025)
by: Su, Rongfeng, et al.
Published: (2025)
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
by: Li, Cancan, et al.
Published: (2025)
by: Li, Cancan, et al.
Published: (2025)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
by: Lin, Chaohao, et al.
Published: (2025)
by: Lin, Chaohao, et al.
Published: (2025)
A Dual-Path Framework with Frequency-and-Time Excited Network for Anomalous Sound Detection
by: Zhang, Yucong, et al.
Published: (2024)
by: Zhang, Yucong, et al.
Published: (2024)
Generalizability of Predictive and Generative Speech Enhancement Models to Pathological Speakers
by: Hou, Mingchi, et al.
Published: (2025)
by: Hou, Mingchi, et al.
Published: (2025)
MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement
by: Yao, Jixun, et al.
Published: (2024)
by: Yao, Jixun, et al.
Published: (2024)
PadAug: Robust Speaker Verification with Simple Waveform-Level Silence Padding
by: Huang, Zijun, et al.
Published: (2025)
by: Huang, Zijun, et al.
Published: (2025)
SpeakerRPL v2: Robust Open-set Speaker Identification through Enhanced Few-shot Foundation Tuning and Model Fusion
by: Chen, Zhiyong, et al.
Published: (2026)
by: Chen, Zhiyong, et al.
Published: (2026)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
by: Guo, Pengcheng, et al.
Published: (2024)
by: Guo, Pengcheng, et al.
Published: (2024)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
by: Chao, Rong, et al.
Published: (2025)
by: Chao, Rong, et al.
Published: (2025)
EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
by: Yao, Jixun, et al.
Published: (2025)
by: Yao, Jixun, et al.
Published: (2025)
Similar Items
-
Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization
by: Tao, Ruijie, et al.
Published: (2024) -
Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System
by: Li, Ze, et al.
Published: (2024) -
The Database and Benchmark for the Source Speaker Tracing Challenge 2024
by: Li, Ze, et al.
Published: (2024) -
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
by: Clarke, Jason, et al.
Published: (2025) -
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
by: Wang, Haoyu, et al.
Published: (2022)