ANIM-400K: A Large-Scale Dataset for Automated End-To-End Dubbing of Video
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Kevin, Liu, Chonghua, Chan, David M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
von: Zhang, Haomin, et al.
Veröffentlicht: (2025)
von: Zhang, Haomin, et al.
Veröffentlicht: (2025)
MCDubber: Multimodal Context-Aware Expressive Video Dubbing
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription
von: Ríos-Vila, Antonio, et al.
Veröffentlicht: (2024)
von: Ríos-Vila, Antonio, et al.
Veröffentlicht: (2024)
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing
von: Liang, Yifan, et al.
Veröffentlicht: (2025)
von: Liang, Yifan, et al.
Veröffentlicht: (2025)
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
von: Zhang, Zhedong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhedong, et al.
Veröffentlicht: (2025)
SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based Approach with Cross-Dataset Evaluation
von: Di Pierno, Andrea, et al.
Veröffentlicht: (2025)
von: Di Pierno, Andrea, et al.
Veröffentlicht: (2025)
Oceanship: A Large-Scale Dataset for Underwater Audio Target Recognition
von: Li, Zeyu, et al.
Veröffentlicht: (2024)
von: Li, Zeyu, et al.
Veröffentlicht: (2024)
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2025)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2025)
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
On the Audio Hallucinations in Large Audio-Video Language Models
von: Nishimura, Taichi, et al.
Veröffentlicht: (2024)
von: Nishimura, Taichi, et al.
Veröffentlicht: (2024)
An End-to-End Speech Summarization Using Large Language Model
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
von: Prabhavalkar, Rohit, et al.
Veröffentlicht: (2024)
von: Prabhavalkar, Rohit, et al.
Veröffentlicht: (2024)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2025)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2025)
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography
von: Dai, Yuqin, et al.
Veröffentlicht: (2025)
von: Dai, Yuqin, et al.
Veröffentlicht: (2025)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
End-to-End Speech-to-Text Translation: A Survey
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
Measuring Sound Symbolism in Audio-visual Models
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2024)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
An investigation of phrase break prediction in an End-to-End TTS system
von: Vadapalli, Anandaswarup
Veröffentlicht: (2023)
von: Vadapalli, Anandaswarup
Veröffentlicht: (2023)
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation
von: Min, Anna, et al.
Veröffentlicht: (2025)
von: Min, Anna, et al.
Veröffentlicht: (2025)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
Training-Free Deepfake Voice Recognition by Leveraging Large-Scale Pre-Trained Models
von: Pianese, Alessandro, et al.
Veröffentlicht: (2024)
von: Pianese, Alessandro, et al.
Veröffentlicht: (2024)
MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
Acoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems
von: Lin, Jhen-Ke, et al.
Veröffentlicht: (2025)
von: Lin, Jhen-Ke, et al.
Veröffentlicht: (2025)
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
von: Le, Trang, et al.
Veröffentlicht: (2024)
von: Le, Trang, et al.
Veröffentlicht: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
von: Moslem, Yasmin
Veröffentlicht: (2024)
von: Moslem, Yasmin
Veröffentlicht: (2024)
SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
von: Wang, Peidong, et al.
Veröffentlicht: (2026)
von: Wang, Peidong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
von: Zhang, Haomin, et al.
Veröffentlicht: (2025) -
MCDubber: Multimodal Context-Aware Expressive Video Dubbing
von: Zhao, Yuan, et al.
Veröffentlicht: (2024) -
Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription
von: Ríos-Vila, Antonio, et al.
Veröffentlicht: (2024) -
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing
von: Liang, Yifan, et al.
Veröffentlicht: (2025) -
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
von: Hu, Rui, et al.
Veröffentlicht: (2025)