Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Hao, Dai, Ju, Zhao, Xin, Zhou, Feng, Pan, Junjun, Li, Lei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Content and Style Aware Audio-Driven Facial Animation
por: Liu, Qingju, et al.
Publicado: (2024)
por: Liu, Qingju, et al.
Publicado: (2024)
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
por: Sun, Zhaokai, et al.
Publicado: (2025)
por: Sun, Zhaokai, et al.
Publicado: (2025)
SemanticAudio: Audio Generation and Editing in Semantic Space
por: Dai, Zheqi, et al.
Publicado: (2026)
por: Dai, Zheqi, et al.
Publicado: (2026)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
por: Guo, Hongming, et al.
Publicado: (2024)
por: Guo, Hongming, et al.
Publicado: (2024)
Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
por: Pan, Zexu, et al.
Publicado: (2025)
por: Pan, Zexu, et al.
Publicado: (2025)
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
por: Chou, Ju-Chieh, et al.
Publicado: (2023)
por: Chou, Ju-Chieh, et al.
Publicado: (2023)
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
por: Kounadis-Bastian, Dionyssos, et al.
Publicado: (2024)
por: Kounadis-Bastian, Dionyssos, et al.
Publicado: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
por: Kwak, Doyeop, et al.
Publicado: (2026)
por: Kwak, Doyeop, et al.
Publicado: (2026)
WavMark: Watermarking for Audio Generation
por: Chen, Guangyu, et al.
Publicado: (2023)
por: Chen, Guangyu, et al.
Publicado: (2023)
PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
por: Shi, Shuchen, et al.
Publicado: (2024)
por: Shi, Shuchen, et al.
Publicado: (2024)
AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound
por: Wijngaard, Gijs, et al.
Publicado: (2025)
por: Wijngaard, Gijs, et al.
Publicado: (2025)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
por: Baser, Oguzhan, et al.
Publicado: (2025)
por: Baser, Oguzhan, et al.
Publicado: (2025)
WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
por: Yuksel, Goksenin, et al.
Publicado: (2025)
por: Yuksel, Goksenin, et al.
Publicado: (2025)
Adapting WavLM for Speech Emotion Recognition
por: Diatlova, Daria, et al.
Publicado: (2024)
por: Diatlova, Daria, et al.
Publicado: (2024)
CoPlay: Audio-agnostic Cognitive Scaling for Acoustic Sensing
por: Li, Yin, et al.
Publicado: (2024)
por: Li, Yin, et al.
Publicado: (2024)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
por: Li, Jiaqi, et al.
Publicado: (2025)
por: Li, Jiaqi, et al.
Publicado: (2025)
Jointly Recognizing Speech and Singing Voices Based on Multi-Task Audio Source Separation
por: Bai, Ye, et al.
Publicado: (2024)
por: Bai, Ye, et al.
Publicado: (2024)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
por: Liu, Qianhui, et al.
Publicado: (2024)
por: Liu, Qianhui, et al.
Publicado: (2024)
LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
por: Chen, Zhipeng, et al.
Publicado: (2026)
por: Chen, Zhipeng, et al.
Publicado: (2026)
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
por: Salganik, Rebecca, et al.
Publicado: (2026)
por: Salganik, Rebecca, et al.
Publicado: (2026)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
por: Xi, Yu, et al.
Publicado: (2024)
por: Xi, Yu, et al.
Publicado: (2024)
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
por: Ma, Hao, et al.
Publicado: (2025)
por: Ma, Hao, et al.
Publicado: (2025)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
por: Hu, Shujie, et al.
Publicado: (2024)
por: Hu, Shujie, et al.
Publicado: (2024)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
por: Li, Jiaqi, et al.
Publicado: (2024)
por: Li, Jiaqi, et al.
Publicado: (2024)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
por: Chen, Yifu, et al.
Publicado: (2025)
por: Chen, Yifu, et al.
Publicado: (2025)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
por: Deng, Keqi, et al.
Publicado: (2024)
por: Deng, Keqi, et al.
Publicado: (2024)
Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech
por: Nakazawa, Kazushi
Publicado: (2026)
por: Nakazawa, Kazushi
Publicado: (2026)
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
por: Li, Xin, et al.
Publicado: (2025)
por: Li, Xin, et al.
Publicado: (2025)
AudioSpa: Spatializing Sound Events with Text
por: Feng, Linfeng, et al.
Publicado: (2025)
por: Feng, Linfeng, et al.
Publicado: (2025)
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
por: Pietroń, Marcin, et al.
Publicado: (2026)
por: Pietroń, Marcin, et al.
Publicado: (2026)
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
por: Li, Yangze, et al.
Publicado: (2024)
por: Li, Yangze, et al.
Publicado: (2024)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
por: Chao, Rong, et al.
Publicado: (2025)
por: Chao, Rong, et al.
Publicado: (2025)
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
por: Stourbe, Theophile, et al.
Publicado: (2024)
por: Stourbe, Theophile, et al.
Publicado: (2024)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
por: Su, Fei, et al.
Publicado: (2026)
por: Su, Fei, et al.
Publicado: (2026)
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition
por: Wang, He, et al.
Publicado: (2024)
por: Wang, He, et al.
Publicado: (2024)
Video-to-Audio Generation with Fine-grained Temporal Semantics
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
por: Kim, Heeseung, et al.
Publicado: (2024)
por: Kim, Heeseung, et al.
Publicado: (2024)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
por: Shan, Weiqiao, et al.
Publicado: (2025)
por: Shan, Weiqiao, et al.
Publicado: (2025)
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
por: Zhao, Xiaohan, et al.
Publicado: (2025)
por: Zhao, Xiaohan, et al.
Publicado: (2025)
A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
por: Sang, Wendi, et al.
Publicado: (2025)
por: Sang, Wendi, et al.
Publicado: (2025)
Ejemplares similares
-
Content and Style Aware Audio-Driven Facial Animation
por: Liu, Qingju, et al.
Publicado: (2024) -
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
por: Sun, Zhaokai, et al.
Publicado: (2025) -
SemanticAudio: Audio Generation and Editing in Semantic Space
por: Dai, Zheqi, et al.
Publicado: (2026) -
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
por: Guo, Hongming, et al.
Publicado: (2024) -
Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
por: Pan, Zexu, et al.
Publicado: (2025)