Masked Contrastive Pre-Training Improves Music Audio Key Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Yonay, Ori, Hammond, Tracy, Yang, Tianbao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
by: Rekimoto, Jun, et al.
Published: (2026)
by: Rekimoto, Jun, et al.
Published: (2026)
Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering
by: Aristorenas, Aris J.
Published: (2024)
by: Aristorenas, Aris J.
Published: (2024)
MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
by: Xu, Tianyu, et al.
Published: (2026)
by: Xu, Tianyu, et al.
Published: (2026)
SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization
by: Mehdi, Naqcho Ali, et al.
Published: (2026)
by: Mehdi, Naqcho Ali, et al.
Published: (2026)
If You Hold Me Without Hurting Me: Pathways to Designing Game Audio for Healthy Escapism and Player Well-being
by: Nunes, Caio, et al.
Published: (2025)
by: Nunes, Caio, et al.
Published: (2025)
Taming Audio VAEs via Target-KL Regularization
by: Seetharaman, Prem, et al.
Published: (2026)
by: Seetharaman, Prem, et al.
Published: (2026)
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
by: Woodard, Brandon, et al.
Published: (2025)
by: Woodard, Brandon, et al.
Published: (2025)
Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis
by: Marmoret, Axel
Published: (2026)
by: Marmoret, Axel
Published: (2026)
Adaptable Symbolic Music Infilling with MIDI-RWKV
by: Zhou-Zheng, Christian, et al.
Published: (2025)
by: Zhou-Zheng, Christian, et al.
Published: (2025)
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
by: Cho, Hyunsung, et al.
Published: (2024)
by: Cho, Hyunsung, et al.
Published: (2024)
Fighting Game Adaptive Background Music for Improved Gameplay
by: Khan, Ibrahim, et al.
Published: (2024)
by: Khan, Ibrahim, et al.
Published: (2024)
Enhanced DareFightingICE Competitions: Sound Design and AI Competitions
by: Khan, Ibrahim, et al.
Published: (2024)
by: Khan, Ibrahim, et al.
Published: (2024)
MaskClip: Detachable Clip-on Piezoelectric Sensing of Mask Surface Vibrations for Real-time Noise-Robust Speech Input
by: Hiraki, Hirotaka, et al.
Published: (2025)
by: Hiraki, Hirotaka, et al.
Published: (2025)
BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps
by: Qian, Lekai, et al.
Published: (2026)
by: Qian, Lekai, et al.
Published: (2026)
Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
by: He, Zhanhong, et al.
Published: (2025)
by: He, Zhanhong, et al.
Published: (2025)
MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation
by: Zhu, Di, et al.
Published: (2026)
by: Zhu, Di, et al.
Published: (2026)
Benchmarking Sub-Genre Classification For Mainstage Dance Music
by: Shu, Hongzhi, et al.
Published: (2024)
by: Shu, Hongzhi, et al.
Published: (2024)
Generation of Musical Timbres using a Text-Guided Diffusion Model
by: Yuan, Weixuan, et al.
Published: (2025)
by: Yuan, Weixuan, et al.
Published: (2025)
Human-Robot Creative Interactions (HRCI): Exploring Creativity in Artificial Agents Using a Story-Telling Game
by: Sandoval, Eduardo Benitez, et al.
Published: (2022)
by: Sandoval, Eduardo Benitez, et al.
Published: (2022)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
by: Wang, Shaowen, et al.
Published: (2025)
by: Wang, Shaowen, et al.
Published: (2025)
OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
by: Mahfi, Muntahi Safwan, et al.
Published: (2025)
by: Mahfi, Muntahi Safwan, et al.
Published: (2025)
Representing Classical Compositions through Implication-Realization Temporal-Gestalt Graphs
by: Bomediano, A. V., et al.
Published: (2025)
by: Bomediano, A. V., et al.
Published: (2025)
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
by: Dhiman, Jai
Published: (2026)
by: Dhiman, Jai
Published: (2026)
Myna: Masking-Based Contrastive Learning of Musical Representations
by: Yonay, Ori, et al.
Published: (2025)
by: Yonay, Ori, et al.
Published: (2025)
Adaptive Background Music for a Fighting Game: A Multi-Instrument Volume Modulation Approach
by: Khan, Ibrahim, et al.
Published: (2023)
by: Khan, Ibrahim, et al.
Published: (2023)
Tactile Melodies: A Desk-Mounted Haptics for Perceiving Musical Experiences
by: Moora, Raj Varshith, et al.
Published: (2024)
by: Moora, Raj Varshith, et al.
Published: (2024)
Sonify Anything: Towards Context-Aware Sonic Interactions in AR
by: Schütz, Laura, et al.
Published: (2025)
by: Schütz, Laura, et al.
Published: (2025)
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
Uncovering Population PK Covariates from VAE-Generated Latent Spaces
by: Perazzolo, Diego, et al.
Published: (2025)
by: Perazzolo, Diego, et al.
Published: (2025)
Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit
by: Soni, Aniket Abhishek
Published: (2025)
by: Soni, Aniket Abhishek
Published: (2025)
Step-Audio-R1 Technical Report
by: Tian, Fei, et al.
Published: (2025)
by: Tian, Fei, et al.
Published: (2025)
Discriminative Subspace Emersion from learning feature relevances across different populations
by: Canducci, Marco, et al.
Published: (2025)
by: Canducci, Marco, et al.
Published: (2025)
Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond
by: Richter-Powell, Jessie, et al.
Published: (2025)
by: Richter-Powell, Jessie, et al.
Published: (2025)
The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS
by: Chen, Kuan-Yu, et al.
Published: (2026)
by: Chen, Kuan-Yu, et al.
Published: (2026)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
by: Khushiyant, et al.
Published: (2026)
by: Khushiyant, et al.
Published: (2026)
Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
BemaGANv2: Discriminator Combination Strategies for GAN-based Vocoders in Long-Term Audio Generation
by: Park, Taesoo, et al.
Published: (2025)
by: Park, Taesoo, et al.
Published: (2025)
Distilled HuBERT for Mobile Speech Emotion Recognition: A Cross-Corpus Validation Study
by: Ismail, Saifelden M.
Published: (2025)
by: Ismail, Saifelden M.
Published: (2025)
SonoHaptics: An Audio-Haptic Cursor for Gaze-Based Object Selection in XR
by: Cho, Hyunsung, et al.
Published: (2024)
by: Cho, Hyunsung, et al.
Published: (2024)
Quantum-Enhanced Analysis and Grading of Vocal Performance
by: Agarwal, Rohan
Published: (2025)
by: Agarwal, Rohan
Published: (2025)
Similar Items
-
NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
by: Rekimoto, Jun, et al.
Published: (2026) -
Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering
by: Aristorenas, Aris J.
Published: (2024) -
MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
by: Xu, Tianyu, et al.
Published: (2026) -
SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization
by: Mehdi, Naqcho Ali, et al.
Published: (2026) -
If You Hold Me Without Hurting Me: Pathways to Designing Game Audio for Healthy Escapism and Player Well-being
by: Nunes, Caio, et al.
Published: (2025)