STONE: Self-supervised Tonality Estimator
Fuente:
arXiv
Salvato in:
| Autori principali: | Kong, Yuexuan, Lostanlen, Vincent, Meseguer-Brocal, Gabriel, Wong, Stella, Lagrange, Mathieu, Hennequin, Romain |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
S-KEY: Self-supervised Learning of Major and Minor Keys from Audio
di: Kong, Yuexuan, et al.
Pubblicazione: (2025)
di: Kong, Yuexuan, et al.
Pubblicazione: (2025)
Emergent musical properties of a transformer under contrastive self-supervised learning
di: Kong, Yuexuan, et al.
Pubblicazione: (2025)
di: Kong, Yuexuan, et al.
Pubblicazione: (2025)
An Experimental Comparison Of Multi-view Self-supervised Methods For Music Tagging
di: Meseguer-Brocal, Gabriel, et al.
Pubblicazione: (2024)
di: Meseguer-Brocal, Gabriel, et al.
Pubblicazione: (2024)
AI-Generated Music Detection and its Challenges
di: Afchar, Darius, et al.
Pubblicazione: (2025)
di: Afchar, Darius, et al.
Pubblicazione: (2025)
Learning to Solve Inverse Problems for Perceptual Sound Matching
di: Han, Han, et al.
Pubblicazione: (2023)
di: Han, Han, et al.
Pubblicazione: (2023)
Detecting music deepfakes is easy but actually hard
di: Afchar, Darius, et al.
Pubblicazione: (2024)
di: Afchar, Darius, et al.
Pubblicazione: (2024)
Multi-Class-Token Transformer for Multitask Self-supervised Music Information Retrieval
di: Kong, Yuexuan, et al.
Pubblicazione: (2025)
di: Kong, Yuexuan, et al.
Pubblicazione: (2025)
STraDa: A Singer Traits Dataset
di: Kong, Yuexuan, et al.
Pubblicazione: (2024)
di: Kong, Yuexuan, et al.
Pubblicazione: (2024)
From Real to Cloned Singer Identification
di: Desblancs, Dorian, et al.
Pubblicazione: (2024)
di: Desblancs, Dorian, et al.
Pubblicazione: (2024)
SCRAPL: Scattering Transform with Random Paths for Machine Learning
di: Mitcheltree, Christopher, et al.
Pubblicazione: (2026)
di: Mitcheltree, Christopher, et al.
Pubblicazione: (2026)
Musical Metamerism with Time--Frequency Scattering
di: Lostanlen, Vincent, et al.
Pubblicazione: (2026)
di: Lostanlen, Vincent, et al.
Pubblicazione: (2026)
Enforcing Speech Content Privacy in Environmental Sound Recordings using Segment-wise Waveform Reversal
di: Tailleur, Modan, et al.
Pubblicazione: (2025)
di: Tailleur, Modan, et al.
Pubblicazione: (2025)
Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
di: Torres, Bernardo, et al.
Pubblicazione: (2025)
di: Torres, Bernardo, et al.
Pubblicazione: (2025)
Mixture of Mixups for Multi-label Classification of Rare Anuran Sounds
di: Moummad, Ilyass, et al.
Pubblicazione: (2024)
di: Moummad, Ilyass, et al.
Pubblicazione: (2024)
Towards better visualizations of urban sound environments: insights from interviews
di: Tailleur, Modan, et al.
Pubblicazione: (2024)
di: Tailleur, Modan, et al.
Pubblicazione: (2024)
Fitting Auditory Filterbanks with Multiresolution Neural Networks
di: Lostanlen, Vincent, et al.
Pubblicazione: (2023)
di: Lostanlen, Vincent, et al.
Pubblicazione: (2023)
Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
di: Frohmann, Markus, et al.
Pubblicazione: (2025)
di: Frohmann, Markus, et al.
Pubblicazione: (2025)
PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
di: Riou, Alain, et al.
Pubblicazione: (2023)
di: Riou, Alain, et al.
Pubblicazione: (2023)
Detection of Deepfake Environmental Audio
di: Ouajdi, Hafsa, et al.
Pubblicazione: (2024)
di: Ouajdi, Hafsa, et al.
Pubblicazione: (2024)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
di: Tao, Dehua, et al.
Pubblicazione: (2024)
di: Tao, Dehua, et al.
Pubblicazione: (2024)
Instabilities in Convnets for Raw Audio
di: Haider, Daniel, et al.
Pubblicazione: (2023)
di: Haider, Daniel, et al.
Pubblicazione: (2023)
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
di: Tailleur, Modan, et al.
Pubblicazione: (2024)
di: Tailleur, Modan, et al.
Pubblicazione: (2024)
Hold Me Tight: Stable Encoder-Decoder Design for Speech Enhancement
di: Haider, Daniel, et al.
Pubblicazione: (2024)
di: Haider, Daniel, et al.
Pubblicazione: (2024)
Probing Self-supervised Learning Models with Target Speech Extraction
di: Peng, Junyi, et al.
Pubblicazione: (2024)
di: Peng, Junyi, et al.
Pubblicazione: (2024)
Self-supervised Reflective Learning through Self-distillation and Online Clustering for Speaker Representation Learning
di: Cai, Danwei, et al.
Pubblicazione: (2024)
di: Cai, Danwei, et al.
Pubblicazione: (2024)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
di: Peng, Junyi, et al.
Pubblicazione: (2024)
di: Peng, Junyi, et al.
Pubblicazione: (2024)
SCDNet: Self-supervised Learning Feature-based Speaker Change Detection
di: Li, Yue, et al.
Pubblicazione: (2024)
di: Li, Yue, et al.
Pubblicazione: (2024)
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
di: Cai, Yiqiang, et al.
Pubblicazione: (2024)
di: Cai, Yiqiang, et al.
Pubblicazione: (2024)
NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
di: Huang, He, et al.
Pubblicazione: (2024)
di: Huang, He, et al.
Pubblicazione: (2024)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024)
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
di: Liang, Xinyu, et al.
Pubblicazione: (2025)
di: Liang, Xinyu, et al.
Pubblicazione: (2025)
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
di: Kuroyanagi, Ibuki, et al.
Pubblicazione: (2025)
di: Kuroyanagi, Ibuki, et al.
Pubblicazione: (2025)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
di: Zaiem, Salah, et al.
Pubblicazione: (2024)
di: Zaiem, Salah, et al.
Pubblicazione: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
di: Gong, Cheng, et al.
Pubblicazione: (2023)
di: Gong, Cheng, et al.
Pubblicazione: (2023)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
Machine listening in a neonatal intensive care unit
di: Tailleur, Modan, et al.
Pubblicazione: (2024)
di: Tailleur, Modan, et al.
Pubblicazione: (2024)
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
di: Premananth, Gowtham, et al.
Pubblicazione: (2024)
di: Premananth, Gowtham, et al.
Pubblicazione: (2024)
ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality
di: Luo, Yu-Xiang, et al.
Pubblicazione: (2025)
di: Luo, Yu-Xiang, et al.
Pubblicazione: (2025)
Performance and energy balance: a comprehensive study of state-of-the-art sound event detection systems
di: Ronchini, Francesca, et al.
Pubblicazione: (2023)
di: Ronchini, Francesca, et al.
Pubblicazione: (2023)
LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations
di: Dutta, Soumya, et al.
Pubblicazione: (2025)
di: Dutta, Soumya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
S-KEY: Self-supervised Learning of Major and Minor Keys from Audio
di: Kong, Yuexuan, et al.
Pubblicazione: (2025) -
Emergent musical properties of a transformer under contrastive self-supervised learning
di: Kong, Yuexuan, et al.
Pubblicazione: (2025) -
An Experimental Comparison Of Multi-view Self-supervised Methods For Music Tagging
di: Meseguer-Brocal, Gabriel, et al.
Pubblicazione: (2024) -
AI-Generated Music Detection and its Challenges
di: Afchar, Darius, et al.
Pubblicazione: (2025) -
Learning to Solve Inverse Problems for Perceptual Sound Matching
di: Han, Han, et al.
Pubblicazione: (2023)