Large-scale unsupervised audio pre-training for video-to-speech synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Kefalas, Triantafyllos, Panagakis, Yannis, Pantic, Maja |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Audio-visual video-to-speech synthesis with synthesized input audio
by: Kefalas, Triantafyllos, et al.
Published: (2023)
by: Kefalas, Triantafyllos, et al.
Published: (2023)
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
by: Smeu, Stefan, et al.
Published: (2024)
by: Smeu, Stefan, et al.
Published: (2024)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2026)
by: Cappellazzo, Umberto, et al.
Published: (2026)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
by: Anand, et al.
Published: (2025)
by: Anand, et al.
Published: (2025)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
Character-aware audio-visual subtitling in context
by: Huh, Jaesung, et al.
Published: (2024)
by: Huh, Jaesung, et al.
Published: (2024)
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
by: Chen, Honglie, et al.
Published: (2024)
by: Chen, Honglie, et al.
Published: (2024)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
by: Cappellazzo, Umberto, et al.
Published: (2024)
by: Cappellazzo, Umberto, et al.
Published: (2024)
Late fusion ensembles for speech recognition on diverse input audio representations
by: Jezidžić, Marin, et al.
Published: (2024)
by: Jezidžić, Marin, et al.
Published: (2024)
Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
by: Maharana, Sarthak Kumar, et al.
Published: (2023)
by: Maharana, Sarthak Kumar, et al.
Published: (2023)
Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol
by: Apostolidis, Konstantinos, et al.
Published: (2024)
by: Apostolidis, Konstantinos, et al.
Published: (2024)
Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
by: Maiti, Soumi, et al.
Published: (2023)
by: Maiti, Soumi, et al.
Published: (2023)
pycnet-audio: A Python package to support bioacoustics data processing
by: Ruff, Zachary J., et al.
Published: (2025)
by: Ruff, Zachary J., et al.
Published: (2025)
When Vision Models Meet Parameter Efficient Look-Aside Adapters Without Large-Scale Audio Pretraining
by: Yeo, Juan, et al.
Published: (2024)
by: Yeo, Juan, et al.
Published: (2024)
Multi-scale Multi-instance Visual Sound Localization and Segmentation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
by: Siuzdak, Hubert
Published: (2023)
by: Siuzdak, Hubert
Published: (2023)
A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation
by: Airale, Louis, et al.
Published: (2023)
by: Airale, Louis, et al.
Published: (2023)
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
by: Ahmed, Tawsif, et al.
Published: (2025)
by: Ahmed, Tawsif, et al.
Published: (2025)
Deep Neural Networks for Automatic Speaker Recognition Do Not Learn Supra-Segmental Temporal Features
by: Neururer, Daniel, et al.
Published: (2023)
by: Neururer, Daniel, et al.
Published: (2023)
Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
by: Bittar, Alexandre, et al.
Published: (2023)
by: Bittar, Alexandre, et al.
Published: (2023)
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
by: Li, Tianqi, et al.
Published: (2024)
by: Li, Tianqi, et al.
Published: (2024)
Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
by: Zhu, Yongxin, et al.
Published: (2024)
by: Zhu, Yongxin, et al.
Published: (2024)
Exploring Federated Self-Supervised Learning for General Purpose Audio Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024)
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024)
Acoustic Scene Classification: A Competition Review
by: Gharib, Shayan, et al.
Published: (2018)
by: Gharib, Shayan, et al.
Published: (2018)
Assessing the Robustness of Spectral Clustering for Deep Speaker Diarization
by: Raghav, Nikhil, et al.
Published: (2024)
by: Raghav, Nikhil, et al.
Published: (2024)
Exploring Green AI for Audio Deepfake Detection
by: Saha, Subhajit, et al.
Published: (2024)
by: Saha, Subhajit, et al.
Published: (2024)
Joint Multimodal Transformer for Emotion Recognition in the Wild
by: Waligora, Paul, et al.
Published: (2024)
by: Waligora, Paul, et al.
Published: (2024)
Dynamic Cross Attention for Audio-Visual Person Verification
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
Dynamic Modality and View Selection for Multimodal Emotion Recognition with Missing Modalities
by: Menon, Luciana Trinkaus, et al.
Published: (2024)
by: Menon, Luciana Trinkaus, et al.
Published: (2024)
An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
by: Malard, Hugo, et al.
Published: (2024)
by: Malard, Hugo, et al.
Published: (2024)
Benchmarking Machine Learning Methods for Distributed Acoustic Sensing
by: Shi, Shuaikai, et al.
Published: (2025)
by: Shi, Shuaikai, et al.
Published: (2025)
Developing an AI-based Integrated System for Bee Health Evaluation
by: Liang, Andrew
Published: (2024)
by: Liang, Andrew
Published: (2024)
Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization
by: Cheng, Luyao, et al.
Published: (2024)
by: Cheng, Luyao, et al.
Published: (2024)
The Solution for Temporal Sound Localisation Task of ICCV 1st Perception Test Challenge 2023
by: Huang, Yurui, et al.
Published: (2024)
by: Huang, Yurui, et al.
Published: (2024)
Towards reliable respiratory disease diagnosis based on cough sounds and vision transformers
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
by: Takahashi, Akira, et al.
Published: (2025)
by: Takahashi, Akira, et al.
Published: (2025)
A High-Accuracy Optical Music Recognition Method Based on Bottleneck Residual Convolutions
by: Ma, Junwen, et al.
Published: (2026)
by: Ma, Junwen, et al.
Published: (2026)
Training-Free Voice Conversion with Factorized Optimal Transport
by: Lobashev, Alexander, et al.
Published: (2025)
by: Lobashev, Alexander, et al.
Published: (2025)
Similar Items
-
Audio-visual video-to-speech synthesis with synthesized input audio
by: Kefalas, Triantafyllos, et al.
Published: (2023) -
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
by: Smeu, Stefan, et al.
Published: (2024) -
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2026) -
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
by: Anand, et al.
Published: (2025) -
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
by: Cappellazzo, Umberto, et al.
Published: (2025)