BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Haliassos, Alexandros, Zinonos, Andreas, Mira, Rodrigo, Petridis, Stavros, Pantic, Maja |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs
by: Haliassos, Alexandros, et al.
Published: (2024)
by: Haliassos, Alexandros, et al.
Published: (2024)
Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition
by: Haliassos, Alexandros, et al.
Published: (2026)
by: Haliassos, Alexandros, et al.
Published: (2026)
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
by: Chen, Honglie, et al.
Published: (2024)
by: Chen, Honglie, et al.
Published: (2024)
FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
by: Zinonos, Andreas, et al.
Published: (2025)
by: Zinonos, Andreas, et al.
Published: (2025)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2026)
by: Cappellazzo, Umberto, et al.
Published: (2026)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
by: Anand, et al.
Published: (2025)
by: Anand, et al.
Published: (2025)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
by: Bigata, Antoni, et al.
Published: (2025)
by: Bigata, Antoni, et al.
Published: (2025)
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
by: Cappellazzo, Umberto, et al.
Published: (2024)
by: Cappellazzo, Umberto, et al.
Published: (2024)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
FaceCrafter: Identity-Conditional Diffusion with Disentangled Control over Facial Pose, Expression, and Emotion
by: Mishima, Kazuaki, et al.
Published: (2025)
by: Mishima, Kazuaki, et al.
Published: (2025)
EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars
by: Drobyshev, Nikita, et al.
Published: (2024)
by: Drobyshev, Nikita, et al.
Published: (2024)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation
by: Bigata, Antoni, et al.
Published: (2025)
by: Bigata, Antoni, et al.
Published: (2025)
Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
by: Gao, Zuan, et al.
Published: (2024)
by: Gao, Zuan, et al.
Published: (2024)
Formula-Supervised Visual-Geometric Pre-training
by: Yamada, Ryosuke, et al.
Published: (2024)
by: Yamada, Ryosuke, et al.
Published: (2024)
In Pursuit of Pixel Supervision for Visual Pre-training
by: Yang, Lihe, et al.
Published: (2025)
by: Yang, Lihe, et al.
Published: (2025)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
by: Yeo, Jeong Hun, et al.
Published: (2025)
by: Yeo, Jeong Hun, et al.
Published: (2025)
Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation
by: Seo, Junyoung, et al.
Published: (2025)
by: Seo, Junyoung, et al.
Published: (2025)
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
by: Kim, Minsu, et al.
Published: (2024)
by: Kim, Minsu, et al.
Published: (2024)
Large-scale unsupervised audio pre-training for video-to-speech synthesis
by: Kefalas, Triantafyllos, et al.
Published: (2023)
by: Kefalas, Triantafyllos, et al.
Published: (2023)
Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
by: Ebouky, Brown, et al.
Published: (2025)
by: Ebouky, Brown, et al.
Published: (2025)
Self-Supervised Pre-Training for Table Structure Recognition Transformer
by: Peng, ShengYun, et al.
Published: (2024)
by: Peng, ShengYun, et al.
Published: (2024)
Effect of Rotation Angle in Self-Supervised Pre-training is Dataset-Dependent
by: Saranchuk, Amy, et al.
Published: (2024)
by: Saranchuk, Amy, et al.
Published: (2024)
A Closer Look at Benchmarking Self-Supervised Pre-training with Image Classification
by: Marks, Markus, et al.
Published: (2024)
by: Marks, Markus, et al.
Published: (2024)
PatchContrast: Self-Supervised Pre-training for 3D Object Detection
by: Shrout, Oren, et al.
Published: (2023)
by: Shrout, Oren, et al.
Published: (2023)
Self-Supervised Pre-training with Combined Datasets for 3D Perception in Autonomous Driving
by: Wang, Shumin, et al.
Published: (2025)
by: Wang, Shumin, et al.
Published: (2025)
Enhancing SAR Object Detection with Self-Supervised Pre-training on Masked Auto-Encoders
by: Pu, Xinyang, et al.
Published: (2025)
by: Pu, Xinyang, et al.
Published: (2025)
GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving
by: Ljungbergh, William, et al.
Published: (2025)
by: Ljungbergh, William, et al.
Published: (2025)
Self-Supervised Pre-training Tasks for an fMRI Time-series Transformer in Autism Detection
by: Zhou, Yinchi, et al.
Published: (2024)
by: Zhou, Yinchi, et al.
Published: (2024)
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
by: Lei, Jiachen, et al.
Published: (2025)
by: Lei, Jiachen, et al.
Published: (2025)
Enhancing Vision-Language Pre-training with Rich Supervisions
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
Improving Pre-trained Self-Supervised Embeddings Through Effective Entropy Maximization
by: Chakraborty, Deep, et al.
Published: (2024)
by: Chakraborty, Deep, et al.
Published: (2024)
Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
PhySU-Net: Long Temporal Context Transformer for rPPG with Self-Supervised Pre-training
by: Savic, Marko, et al.
Published: (2024)
by: Savic, Marko, et al.
Published: (2024)
VILA: On Pre-training for Visual Language Models
by: Lin, Ji, et al.
Published: (2023)
by: Lin, Ji, et al.
Published: (2023)
Endo-CLIP: Progressive Self-Supervised Pre-training on Raw Colonoscopy Records
by: He, Yili, et al.
Published: (2025)
by: He, Yili, et al.
Published: (2025)
E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
by: Zhao, Qitao, et al.
Published: (2025)
by: Zhao, Qitao, et al.
Published: (2025)
Similar Items
-
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs
by: Haliassos, Alexandros, et al.
Published: (2024) -
Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition
by: Haliassos, Alexandros, et al.
Published: (2026) -
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
by: Chen, Honglie, et al.
Published: (2024) -
FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
by: Zinonos, Andreas, et al.
Published: (2025) -
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2026)