AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Jeongsoo, Park, Se Jin, Kim, Minsu, Ro, Yong Man |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation
von: Cho, Sungwoo, et al.
Veröffentlicht: (2025)
von: Cho, Sungwoo, et al.
Veröffentlicht: (2025)
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models
von: Tseng, Yuan, et al.
Veröffentlicht: (2023)
von: Tseng, Yuan, et al.
Veröffentlicht: (2023)
Text-driven Talking Face Synthesis by Reprogramming Audio-driven Models
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification
von: Xiao, Zhenyuan, et al.
Veröffentlicht: (2024)
von: Xiao, Zhenyuan, et al.
Veröffentlicht: (2024)
Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
von: Li, Kai, et al.
Veröffentlicht: (2023)
von: Li, Kai, et al.
Veröffentlicht: (2023)
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2024)
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2024)
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
von: Chen, Honglie, et al.
Veröffentlicht: (2024)
von: Chen, Honglie, et al.
Veröffentlicht: (2024)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
von: Park, Jisoo, et al.
Veröffentlicht: (2025)
von: Park, Jisoo, et al.
Veröffentlicht: (2025)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
Getting More for Less: Using Weak Labels and AV-Mixup for Robust Audio-Visual Speaker Verification
von: Selvakumar, Anith, et al.
Veröffentlicht: (2023)
von: Selvakumar, Anith, et al.
Veröffentlicht: (2023)
From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation
von: Su, Kun, et al.
Veröffentlicht: (2024)
von: Su, Kun, et al.
Veröffentlicht: (2024)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion
von: Lim, DongHoon, et al.
Veröffentlicht: (2025)
von: Lim, DongHoon, et al.
Veröffentlicht: (2025)
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
von: Nakada, Shota, et al.
Veröffentlicht: (2024)
von: Nakada, Shota, et al.
Veröffentlicht: (2024)
3D Audio-Visual Segmentation
von: Sokolov, Artem, et al.
Veröffentlicht: (2024)
von: Sokolov, Artem, et al.
Veröffentlicht: (2024)
SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model
von: Qian, Xinyuan, et al.
Veröffentlicht: (2024)
von: Qian, Xinyuan, et al.
Veröffentlicht: (2024)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
von: Xu, Le, et al.
Veröffentlicht: (2025)
von: Xu, Le, et al.
Veröffentlicht: (2025)
Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
von: Gong, Kaixiong, et al.
Veröffentlicht: (2024)
von: Gong, Kaixiong, et al.
Veröffentlicht: (2024)
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
von: Li, Cancan, et al.
Veröffentlicht: (2025)
von: Li, Cancan, et al.
Veröffentlicht: (2025)
Audio-Visual Speech Separation via Bottleneck Iterative Network
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
von: Yang, Shiqi, et al.
Veröffentlicht: (2024)
von: Yang, Shiqi, et al.
Veröffentlicht: (2024)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition
von: Dai, Yusheng, et al.
Veröffentlicht: (2024)
von: Dai, Yusheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023) -
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025) -
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025) -
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025) -
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)