Remastering Divide and Remaster: A Cinematic Audio Source Separation Dataset with Multilingual Support
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Watcharasupat, Karn N., Wu, Chih-Wei, Orife, Iroro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
Zero-Shot Crate Digging: DJ Tool Retrieval Using Speech Activity, Music Structure And CLAP Embeddings
von: Orife, Iroro
Veröffentlicht: (2024)
von: Orife, Iroro
Veröffentlicht: (2024)
Separate This, and All of these Things Around It: Music Source Separation via Hyperellipsoidal Queries
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2025)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2025)
Looking Similar, Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation Learning
von: Singh, Nikhil, et al.
Veröffentlicht: (2023)
von: Singh, Nikhil, et al.
Veröffentlicht: (2023)
A Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four Stems
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
Quantifying Spatial Audio Quality Impairment
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus
von: Ogunremi, Tolulope, et al.
Veröffentlicht: (2023)
von: Ogunremi, Tolulope, et al.
Veröffentlicht: (2023)
DnR-nonverbal: Cinematic Audio Source Separation Dataset Containing Non-Verbal Sounds
von: Hasumi, Takuya, et al.
Veröffentlicht: (2025)
von: Hasumi, Takuya, et al.
Veröffentlicht: (2025)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
Uncertainty Estimation in the Real World: A Study on Music Emotion Recognition
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2025)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2025)
Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025)
Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR
von: Bai, Junwen, et al.
Veröffentlicht: (2024)
von: Bai, Junwen, et al.
Veröffentlicht: (2024)
Anatomy of Industrial Scale Multilingual ASR
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024)
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024)
FRCRN: Boosting Feature Representation using Frequency Recurrence for Monaural Speech Enhancement
von: Zhao, Shengkui, et al.
Veröffentlicht: (2022)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2022)
Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark
von: Ciobanu, Ioan-Paul, et al.
Veröffentlicht: (2025)
von: Ciobanu, Ioan-Paul, et al.
Veröffentlicht: (2025)
Improving Text-To-Audio Models with Synthetic Captions
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
von: Beyene, Luel Hagos, et al.
Veröffentlicht: (2025)
von: Beyene, Luel Hagos, et al.
Veröffentlicht: (2025)
SynthSOD: Developing an Heterogeneous Dataset for Orchestra Music Source Separation
von: Garcia-Martinez, Jaime, et al.
Veröffentlicht: (2024)
von: Garcia-Martinez, Jaime, et al.
Veröffentlicht: (2024)
ETTA: Elucidating the Design Space of Text-to-Audio Models
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Cascaded Cross-Modal Transformer for Audio-Textual Classification
von: Ristea, Nicolae-Catalin, et al.
Veröffentlicht: (2024)
von: Ristea, Nicolae-Catalin, et al.
Veröffentlicht: (2024)
Audio-to-Score Conversion Model Based on Whisper methodology
von: Zhang, Hongyao, et al.
Veröffentlicht: (2024)
von: Zhang, Hongyao, et al.
Veröffentlicht: (2024)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
tinyCLAP: Distilling Constrastive Language-Audio Pretrained Models
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
Classification of Spontaneous and Scripted Speech for Multilingual Audio
von: Elisha, Shahar, et al.
Veröffentlicht: (2024)
von: Elisha, Shahar, et al.
Veröffentlicht: (2024)
Papez: Resource-Efficient Speech Separation with Auditory Working Memory
von: Oh, Hyunseok, et al.
Veröffentlicht: (2024)
von: Oh, Hyunseok, et al.
Veröffentlicht: (2024)
ARAUS: A Large-Scale Dataset and Baseline Models of Affective Responses to Augmented Urban Soundscapes
von: Ooi, Kenneth, et al.
Veröffentlicht: (2022)
von: Ooi, Kenneth, et al.
Veröffentlicht: (2022)
Multiple-Instance, Cascaded Classification for Keyword Spotting in Narrow-Band Audio
von: AbdulKader, Ahmad, et al.
Veröffentlicht: (2017)
von: AbdulKader, Ahmad, et al.
Veröffentlicht: (2017)
The Spheres Dataset: Multitrack Orchestral Recordings for Music Source Separation and Information Retrieval
von: Garcia-Martinez, Jaime, et al.
Veröffentlicht: (2025)
von: Garcia-Martinez, Jaime, et al.
Veröffentlicht: (2025)
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
von: Caubrière, Antoine, et al.
Veröffentlicht: (2024)
von: Caubrière, Antoine, et al.
Veröffentlicht: (2024)
XLS-R Deep Learning Model for Multilingual ASR on Low- Resource Languages: Indonesian, Javanese, and Sundanese
von: Arisaputra, Panji, et al.
Veröffentlicht: (2024)
von: Arisaputra, Panji, et al.
Veröffentlicht: (2024)
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
von: Amooie, Reihaneh, et al.
Veröffentlicht: (2025)
von: Amooie, Reihaneh, et al.
Veröffentlicht: (2025)
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
von: Vieting, Peter, et al.
Veröffentlicht: (2023)
von: Vieting, Peter, et al.
Veröffentlicht: (2023)
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
von: Yan, Brian, et al.
Veröffentlicht: (2025)
von: Yan, Brian, et al.
Veröffentlicht: (2025)
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
von: Papi, Sara, et al.
Veröffentlicht: (2023)
von: Papi, Sara, et al.
Veröffentlicht: (2023)
Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
von: Choi, Youngwon, et al.
Veröffentlicht: (2025)
von: Choi, Youngwon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024) -
A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023) -
Zero-Shot Crate Digging: DJ Tool Retrieval Using Speech Activity, Music Structure And CLAP Embeddings
von: Orife, Iroro
Veröffentlicht: (2024) -
Separate This, and All of these Things Around It: Music Source Separation via Hyperellipsoidal Queries
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2025) -
Looking Similar, Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation Learning
von: Singh, Nikhil, et al.
Veröffentlicht: (2023)