A Comprehensive Review and Taxonomy of Audio-Visual Synchronization Techniques for Realistic Speech Animation
Fuente:
arXiv
Guardado en:
| Autores principales: | Fernandes, Jose Geraldo, Nascimento, Sinval, Dominguete, Daniel, Oliveira, André, Rotsen, Lucas, Souza, Gabriel, Brochero, David, Facury, Luiz, Vilela, Mateus, Costa, Hebert, Coelho, Frederico, Braga, Antônio P. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Audio-Visual Feature Synchronization for Robust Speech Enhancement in Hearing Aids
por: Saleem, Nasir, et al.
Publicado: (2025)
por: Saleem, Nasir, et al.
Publicado: (2025)
Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars
por: NVIDIA, et al.
Publicado: (2025)
por: NVIDIA, et al.
Publicado: (2025)
Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems
por: Xiao, Yang, et al.
Publicado: (2026)
por: Xiao, Yang, et al.
Publicado: (2026)
Silent Speech Interfaces in the Era of Large Language Models: A Comprehensive Taxonomy and Systematic Review
por: Xu, Kele, et al.
Publicado: (2026)
por: Xu, Kele, et al.
Publicado: (2026)
Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech
por: O'Reilly, Patrick, et al.
Publicado: (2025)
por: O'Reilly, Patrick, et al.
Publicado: (2025)
SALT: Standardized Audio event Label Taxonomy
por: Stamatiadis, Paraskevas, et al.
Publicado: (2024)
por: Stamatiadis, Paraskevas, et al.
Publicado: (2024)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
por: Shi, Jiatong, et al.
Publicado: (2024)
por: Shi, Jiatong, et al.
Publicado: (2024)
Content and Style Aware Audio-Driven Facial Animation
por: Liu, Qingju, et al.
Publicado: (2024)
por: Liu, Qingju, et al.
Publicado: (2024)
Meta-Learning in Audio and Speech Processing: An End to End Comprehensive Review
por: Raimon, Athul, et al.
Publicado: (2024)
por: Raimon, Athul, et al.
Publicado: (2024)
Application of Audio Fingerprinting Techniques for Real-Time Scalable Speech Retrieval and Speech Clusterization
por: Altwlkany, Kemal, et al.
Publicado: (2024)
por: Altwlkany, Kemal, et al.
Publicado: (2024)
The Artist is Present: Traces of Artists Resigind and Spawning in Text-to-Audio AI
por: Coelho, Guilherme
Publicado: (2025)
por: Coelho, Guilherme
Publicado: (2025)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
por: Wang, Le, et al.
Publicado: (2025)
por: Wang, Le, et al.
Publicado: (2025)
Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation
por: Li, Hao, et al.
Publicado: (2025)
por: Li, Hao, et al.
Publicado: (2025)
Semantic and Semiotic Interplays in Text-to-Audio AI: Exploring Cognitive Dynamics and Musical Interactions
por: Coelho, Guilherme
Publicado: (2025)
por: Coelho, Guilherme
Publicado: (2025)
Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
por: Chen, Xuanjun, et al.
Publicado: (2025)
por: Chen, Xuanjun, et al.
Publicado: (2025)
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
por: Li, Baihan, et al.
Publicado: (2024)
por: Li, Baihan, et al.
Publicado: (2024)
Audio-Visual Speech Enhancement for Spatial Audio - Spatial-VisualVoice and the MAVE Database
por: Yaffe, Danielle, et al.
Publicado: (2025)
por: Yaffe, Danielle, et al.
Publicado: (2025)
Interpreting the Role of Visemes in Audio-Visual Speech Recognition
por: Papadopoulos, Aristeidis, et al.
Publicado: (2025)
por: Papadopoulos, Aristeidis, et al.
Publicado: (2025)
Text-Independent Speaker Identification Using Audio Looping With Margin Based Loss Functions
por: Garcia, Elliot Q C, et al.
Publicado: (2025)
por: Garcia, Elliot Q C, et al.
Publicado: (2025)
Enhancing Crowdsourced Audio for Text-to-Speech Models
por: Giraldo, José, et al.
Publicado: (2024)
por: Giraldo, José, et al.
Publicado: (2024)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
por: Li, Jiaqi, et al.
Publicado: (2024)
por: Li, Jiaqi, et al.
Publicado: (2024)
Speech Separation using Neural Audio Codecs with Embedding Loss
por: Yip, Jia Qi, et al.
Publicado: (2024)
por: Yip, Jia Qi, et al.
Publicado: (2024)
ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals
por: E, Ameenudeen P, et al.
Publicado: (2026)
por: E, Ameenudeen P, et al.
Publicado: (2026)
Multi-Speaker Conversational Audio Deepfake: Taxonomy, Dataset and Pilot Study
por: Ahmed, Alabi, et al.
Publicado: (2026)
por: Ahmed, Alabi, et al.
Publicado: (2026)
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
por: de Seyssel, Maureen, et al.
Publicado: (2025)
por: de Seyssel, Maureen, et al.
Publicado: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
por: Zhou, Dongliang, et al.
Publicado: (2025)
por: Zhou, Dongliang, et al.
Publicado: (2025)
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
por: Deng, Keqi, et al.
Publicado: (2023)
por: Deng, Keqi, et al.
Publicado: (2023)
Hello-Chat: Towards Realistic Social Audio Interactions
por: Hou, Yueran, et al.
Publicado: (2026)
por: Hou, Yueran, et al.
Publicado: (2026)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
por: Lin, Zhaofeng, et al.
Publicado: (2024)
por: Lin, Zhaofeng, et al.
Publicado: (2024)
CUSIDE-array: A Streaming Multi-Channel End-to-End Speech Recognition System with Realistic Evaluations
por: Kong, Xiangzhu, et al.
Publicado: (2024)
por: Kong, Xiangzhu, et al.
Publicado: (2024)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
por: Kim, Jongsuk, et al.
Publicado: (2025)
por: Kim, Jongsuk, et al.
Publicado: (2025)
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
por: Yang, Xiaoyu, et al.
Publicado: (2025)
por: Yang, Xiaoyu, et al.
Publicado: (2025)
VoxEffects: A Speech-Oriented Audio Effects Dataset and Benchmark
por: Zhang, Zhe, et al.
Publicado: (2026)
por: Zhang, Zhe, et al.
Publicado: (2026)
Harmonic Detection from Noisy Speech with Auditory Frame Gain for Intelligibility Enhancement
por: Queiroz, A., et al.
Publicado: (2024)
por: Queiroz, A., et al.
Publicado: (2024)
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
por: Zhang, Haoyang, et al.
Publicado: (2026)
por: Zhang, Haoyang, et al.
Publicado: (2026)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
por: Lei, Ke, et al.
Publicado: (2026)
por: Lei, Ke, et al.
Publicado: (2026)
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
por: Zhao, Xiaohan, et al.
Publicado: (2025)
por: Zhao, Xiaohan, et al.
Publicado: (2025)
DASB - Discrete Audio and Speech Benchmark
por: Mousavi, Pooneh, et al.
Publicado: (2024)
por: Mousavi, Pooneh, et al.
Publicado: (2024)
SynHate: Detecting Hate Speech in Synthetic Deepfake Audio
por: Ranjan, Rishabh, et al.
Publicado: (2025)
por: Ranjan, Rishabh, et al.
Publicado: (2025)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
por: Zhang, Xueyao, et al.
Publicado: (2023)
por: Zhang, Xueyao, et al.
Publicado: (2023)
Ejemplares similares
-
Audio-Visual Feature Synchronization for Robust Speech Enhancement in Hearing Aids
por: Saleem, Nasir, et al.
Publicado: (2025) -
Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars
por: NVIDIA, et al.
Publicado: (2025) -
Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems
por: Xiao, Yang, et al.
Publicado: (2026) -
Silent Speech Interfaces in the Era of Large Language Models: A Comprehensive Taxonomy and Systematic Review
por: Xu, Kele, et al.
Publicado: (2026) -
Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech
por: O'Reilly, Patrick, et al.
Publicado: (2025)