Cluster and Separate: a GNN Approach to Voice and Staff Prediction for Score Engraving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Foscarin, Francesco, Karystinaios, Emmanouil, Nakamura, Eita, Widmer, Gerhard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Perception-Inspired Graph Convolution for Music Understanding Tasks
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024)
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024)
SMUG-Explain: A Framework for Symbolic Music Graph Explanations
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024)
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024)
EngravingGNN: A Hybrid Graph Neural Network for End-to-End Piano Score Engraving
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2025)
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2025)
GraphMuse: A Library for Symbolic Music Graph Processing
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024)
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024)
Multi-Stage Music Source Restoration with BandSplit-RoFormer Separation and HiFi++ GAN
von: Morocutti, Tobias, et al.
Veröffentlicht: (2026)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2026)
Beat this! Accurate beat tracking without DBN postprocessing
von: Foscarin, Francesco, et al.
Veröffentlicht: (2024)
von: Foscarin, Francesco, et al.
Veröffentlicht: (2024)
Exploring Performance-Complexity Trade-Offs in Sound Event Detection Models
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
On Temporal Guidance and Iterative Refinement in Audio Source Separation
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation
von: Karystinaios, Emmanouil
Veröffentlicht: (2025)
von: Karystinaios, Emmanouil
Veröffentlicht: (2025)
Device-Robust Acoustic Scene Classification via Impulse Response Augmentation
von: Morocutti, Tobias, et al.
Veröffentlicht: (2023)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2023)
Language Models for Music Medicine Generation
von: Nikolakakis, Emmanouil, et al.
Veröffentlicht: (2024)
von: Nikolakakis, Emmanouil, et al.
Veröffentlicht: (2024)
A Knowledge-Driven Approach to Music Segmentation, Music Source Separation and Cinematic Audio Source Separation
von: Ho, Chun-wei, et al.
Veröffentlicht: (2026)
von: Ho, Chun-wei, et al.
Veröffentlicht: (2026)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
Sounding Out Reconstruction Error-Based Evaluation of Generative Models of Expressive Performance
von: Peter, Silvan David, et al.
Veröffentlicht: (2023)
von: Peter, Silvan David, et al.
Veröffentlicht: (2023)
AFEN: Respiratory Disease Classification using Ensemble Learning
von: Nadkarni, Rahul, et al.
Veröffentlicht: (2024)
von: Nadkarni, Rahul, et al.
Veröffentlicht: (2024)
MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2026)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2026)
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
Mitigating Unauthorized Speech Synthesis for Voice Protection
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2024)
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
von: Ma, Yinghao, et al.
Veröffentlicht: (2025)
von: Ma, Yinghao, et al.
Veröffentlicht: (2025)
The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
von: Baba, Kaito, et al.
Veröffentlicht: (2024)
von: Baba, Kaito, et al.
Veröffentlicht: (2024)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
von: Nejad, Mahsa Ghazvini, et al.
Veröffentlicht: (2025)
von: Nejad, Mahsa Ghazvini, et al.
Veröffentlicht: (2025)
CoMoSVC: Consistency Model-based Singing Voice Conversion
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
Personalized Speech Enhancement Without a Separate Speaker Embedding Model
von: Pärnamaa, Tanel, et al.
Veröffentlicht: (2024)
von: Pärnamaa, Tanel, et al.
Veröffentlicht: (2024)
A Conditioned UNet for Music Source Separation
von: O'Hanlon, Ken, et al.
Veröffentlicht: (2025)
von: O'Hanlon, Ken, et al.
Veröffentlicht: (2025)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
Blind Separation of Vibration Sources using Deep Learning and Deconvolution
von: Makienko, Igor, et al.
Veröffentlicht: (2024)
von: Makienko, Igor, et al.
Veröffentlicht: (2024)
Reproducible Machine Learning-based Voice Pathology Detection: Introducing the Pitch Difference Feature
von: Vrba, Jan, et al.
Veröffentlicht: (2024)
von: Vrba, Jan, et al.
Veröffentlicht: (2024)
Training-Free Multi-Step Audio Source Separation
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
Text-Queried Audio Source Separation via Hierarchical Modeling
von: Yin, Xinlei, et al.
Veröffentlicht: (2025)
von: Yin, Xinlei, et al.
Veröffentlicht: (2025)
What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
A Data-Driven Diffusion-based Approach for Audio Deepfake Explanations
von: Grinberg, Petr, et al.
Veröffentlicht: (2025)
von: Grinberg, Petr, et al.
Veröffentlicht: (2025)
Neural Blind Source Separation and Diarization for Distant Speech Recognition
von: Bando, Yoshiaki, et al.
Veröffentlicht: (2024)
von: Bando, Yoshiaki, et al.
Veröffentlicht: (2024)
Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation
von: Chen, Ke, et al.
Veröffentlicht: (2024)
von: Chen, Ke, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Perception-Inspired Graph Convolution for Music Understanding Tasks
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024) -
SMUG-Explain: A Framework for Symbolic Music Graph Explanations
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024) -
EngravingGNN: A Hybrid Graph Neural Network for End-to-End Piano Score Engraving
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2025) -
GraphMuse: A Library for Symbolic Music Graph Processing
von: Karystinaios, Emmanouil, et al.
Veröffentlicht: (2024) -
Multi-Stage Music Source Restoration with BandSplit-RoFormer Separation and HiFi++ GAN
von: Morocutti, Tobias, et al.
Veröffentlicht: (2026)