Salvato in:
| Autori principali: | Sabra, Adam, Wronka, Cyprian, Mao, Michelle, Hijazi, Samer |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2402.12482 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MUSE: Flexible Voiceprint Receptive Fields and Multi-Path Fusion Enhanced Taylor Transformer for U-Net-based Speech Enhancement
di: Lin, Zizhen, et al.
Pubblicazione: (2024)
di: Lin, Zizhen, et al.
Pubblicazione: (2024)
Application of Audio Fingerprinting Techniques for Real-Time Scalable Speech Retrieval and Speech Clusterization
di: Altwlkany, Kemal, et al.
Pubblicazione: (2024)
di: Altwlkany, Kemal, et al.
Pubblicazione: (2024)
Analyzing and reducing the synthetic-to-real transfer gap in Music Information Retrieval: the task of automatic drum transcription
di: Zehren, Mickaël, et al.
Pubblicazione: (2024)
di: Zehren, Mickaël, et al.
Pubblicazione: (2024)
Transfer Learning with Semi-Supervised Dataset Annotation for Birdcall Classification
di: Miyaguchi, Anthony, et al.
Pubblicazione: (2023)
di: Miyaguchi, Anthony, et al.
Pubblicazione: (2023)
Music Foundation Model as Generic Booster for Music Downstream Tasks
di: Liao, WeiHsiang, et al.
Pubblicazione: (2024)
di: Liao, WeiHsiang, et al.
Pubblicazione: (2024)
Hybrid Losses for Hierarchical Embedding Learning
di: Tian, Haokun, et al.
Pubblicazione: (2025)
di: Tian, Haokun, et al.
Pubblicazione: (2025)
Deconstructing Jazz Piano Style Using Machine Learning
di: Cheston, Huw, et al.
Pubblicazione: (2025)
di: Cheston, Huw, et al.
Pubblicazione: (2025)
A Novel Audio Representation for Music Genre Identification in MIR
di: Kamuni, Navin, et al.
Pubblicazione: (2024)
di: Kamuni, Navin, et al.
Pubblicazione: (2024)
EAViT: External Attention Vision Transformer for Audio Classification
di: Iqbal, Aquib, et al.
Pubblicazione: (2024)
di: Iqbal, Aquib, et al.
Pubblicazione: (2024)
Nested Music Transformer: Sequentially Decoding Compound Tokens in Symbolic Music and Audio Generation
di: Yoo, HaeJun, et al.
Pubblicazione: (2024)
di: Yoo, HaeJun, et al.
Pubblicazione: (2024)
Multi-label Cross-lingual automatic music genre classification from lyrics with Sentence BERT
di: Tavares, Tiago Fernandes, et al.
Pubblicazione: (2025)
di: Tavares, Tiago Fernandes, et al.
Pubblicazione: (2025)
Improving Musical Instrument Classification with Advanced Machine Learning Techniques
di: Chulev, Joanikij
Pubblicazione: (2024)
di: Chulev, Joanikij
Pubblicazione: (2024)
Dissecting Temporal Understanding in Text-to-Audio Retrieval
di: Oncescu, Andreea-Maria, et al.
Pubblicazione: (2024)
di: Oncescu, Andreea-Maria, et al.
Pubblicazione: (2024)
Emergent musical properties of a transformer under contrastive self-supervised learning
di: Kong, Yuexuan, et al.
Pubblicazione: (2025)
di: Kong, Yuexuan, et al.
Pubblicazione: (2025)
Uncertainty Estimation in the Real World: A Study on Music Emotion Recognition
di: Watcharasupat, Karn N., et al.
Pubblicazione: (2025)
di: Watcharasupat, Karn N., et al.
Pubblicazione: (2025)
Separate This, and All of these Things Around It: Music Source Separation via Hyperellipsoidal Queries
di: Watcharasupat, Karn N., et al.
Pubblicazione: (2025)
di: Watcharasupat, Karn N., et al.
Pubblicazione: (2025)
Universal Music Representations? Evaluating Foundation Models on World Music Corpora
di: Papaioannou, Charilaos, et al.
Pubblicazione: (2025)
di: Papaioannou, Charilaos, et al.
Pubblicazione: (2025)
From Real to Cloned Singer Identification
di: Desblancs, Dorian, et al.
Pubblicazione: (2024)
di: Desblancs, Dorian, et al.
Pubblicazione: (2024)
Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness
di: Loweimi, Erfan, et al.
Pubblicazione: (2025)
di: Loweimi, Erfan, et al.
Pubblicazione: (2025)
Adaptive Slimming for Scalable and Efficient Speech Enhancement
di: Miccini, Riccardo, et al.
Pubblicazione: (2025)
di: Miccini, Riccardo, et al.
Pubblicazione: (2025)
Scalable Speech Enhancement with Dynamic Channel Pruning
di: Miccini, Riccardo, et al.
Pubblicazione: (2024)
di: Miccini, Riccardo, et al.
Pubblicazione: (2024)
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech
di: de Oliveira, Danilo, et al.
Pubblicazione: (2024)
di: de Oliveira, Danilo, et al.
Pubblicazione: (2024)
MERGE -- A Bimodal Audio-Lyrics Dataset for Static Music Emotion Recognition
di: Louro, Pedro Lima, et al.
Pubblicazione: (2024)
di: Louro, Pedro Lima, et al.
Pubblicazione: (2024)
On the Effect of Data-Augmentation on Local Embedding Properties in the Contrastive Learning of Music Audio Representations
di: McCallum, Matthew C., et al.
Pubblicazione: (2024)
di: McCallum, Matthew C., et al.
Pubblicazione: (2024)
A Dataset and Baselines for Measuring and Predicting the Music Piece Memorability
di: Tseng, Li-Yang, et al.
Pubblicazione: (2024)
di: Tseng, Li-Yang, et al.
Pubblicazione: (2024)
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
di: Fu, Siyuan, et al.
Pubblicazione: (2025)
di: Fu, Siyuan, et al.
Pubblicazione: (2025)
Music Genre Classification: Ensemble Learning with Subcomponents-level Attention
di: Liu, Yichen, et al.
Pubblicazione: (2024)
di: Liu, Yichen, et al.
Pubblicazione: (2024)
Learning Normal Patterns in Musical Loops
di: Dadman, Shayan, et al.
Pubblicazione: (2025)
di: Dadman, Shayan, et al.
Pubblicazione: (2025)
Equivariance-based self-supervised learning for audio signal recovery from clipped measurements
di: Sechaud, Victor, et al.
Pubblicazione: (2024)
di: Sechaud, Victor, et al.
Pubblicazione: (2024)
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
di: Abootorabi, Mohammad Mahdi, et al.
Pubblicazione: (2024)
di: Abootorabi, Mohammad Mahdi, et al.
Pubblicazione: (2024)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
di: Lin, Chyi-Jiunn, et al.
Pubblicazione: (2024)
di: Lin, Chyi-Jiunn, et al.
Pubblicazione: (2024)
EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
di: Richter, Julius, et al.
Pubblicazione: (2024)
di: Richter, Julius, et al.
Pubblicazione: (2024)
Investigating the Effects of Diffusion-based Conditional Generative Speech Models Used for Speech Enhancement on Dysarthric Speech
di: Reszka, Joanna, et al.
Pubblicazione: (2024)
di: Reszka, Joanna, et al.
Pubblicazione: (2024)
Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement
di: Sadeghi, Mostafa, et al.
Pubblicazione: (2025)
di: Sadeghi, Mostafa, et al.
Pubblicazione: (2025)
Test-Time Training for Speech Enhancement
di: Behera, Avishkar, et al.
Pubblicazione: (2025)
di: Behera, Avishkar, et al.
Pubblicazione: (2025)
Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance
di: Gong, Ziling, et al.
Pubblicazione: (2026)
di: Gong, Ziling, et al.
Pubblicazione: (2026)
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
di: Weck, Benno, et al.
Pubblicazione: (2023)
di: Weck, Benno, et al.
Pubblicazione: (2023)
Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for Tempo Prediction and Search
di: McCallum, Matthew C., et al.
Pubblicazione: (2024)
di: McCallum, Matthew C., et al.
Pubblicazione: (2024)
Documenti analoghi
-
MUSE: Flexible Voiceprint Receptive Fields and Multi-Path Fusion Enhanced Taylor Transformer for U-Net-based Speech Enhancement
di: Lin, Zizhen, et al.
Pubblicazione: (2024) -
Application of Audio Fingerprinting Techniques for Real-Time Scalable Speech Retrieval and Speech Clusterization
di: Altwlkany, Kemal, et al.
Pubblicazione: (2024) -
Analyzing and reducing the synthetic-to-real transfer gap in Music Information Retrieval: the task of automatic drum transcription
di: Zehren, Mickaël, et al.
Pubblicazione: (2024) -
Transfer Learning with Semi-Supervised Dataset Annotation for Birdcall Classification
di: Miyaguchi, Anthony, et al.
Pubblicazione: (2023) -
Music Foundation Model as Generic Booster for Music Downstream Tasks
di: Liao, WeiHsiang, et al.
Pubblicazione: (2024)