Filling MIDI Velocity using U-Net Image Colorizer
Fuente:
arXiv
Guardado en:
| Autores principales: | He, Zhanhong, Cooper, David, Huang, Defeng, Togneri, Roberto |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Score-Informed Transformer for Refining MIDI Velocity in Automatic Music Transcription
por: He, Zhanhong, et al.
Publicado: (2025)
por: He, Zhanhong, et al.
Publicado: (2025)
How to Infer Repeat Structures in MIDI Performances
por: Peter, Silvan, et al.
Publicado: (2025)
por: Peter, Silvan, et al.
Publicado: (2025)
Pseudo Strong Labels from Frame-Level Predictions for Weakly Supervised Sound Event Detection
por: Zhang, Yuliang, et al.
Publicado: (2025)
por: Zhang, Yuliang, et al.
Publicado: (2025)
Impact of Noisy Labels on Sound Event Detection: Deletion Errors Are More Detrimental Than Insertion Errors
por: Zhang, Yuliang, et al.
Publicado: (2024)
por: Zhang, Yuliang, et al.
Publicado: (2024)
MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline
por: Tsai, Fang-Duo, et al.
Publicado: (2026)
por: Tsai, Fang-Duo, et al.
Publicado: (2026)
Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
por: He, Zhanhong, et al.
Publicado: (2025)
por: He, Zhanhong, et al.
Publicado: (2025)
Zero to 16383 Through the Wire: Transmitting High- Resolution MIDI with WebSockets and the Browser
por: McKemie, Daniel
Publicado: (2025)
por: McKemie, Daniel
Publicado: (2025)
Dance2MIDI: Dance-driven multi-instruments music generation
por: Han, Bo, et al.
Publicado: (2023)
por: Han, Bo, et al.
Publicado: (2023)
Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations
por: Wachter, Maximilian, et al.
Publicado: (2026)
por: Wachter, Maximilian, et al.
Publicado: (2026)
Expressive MIDI-format Piano Performance Generation
por: Liu, Jingwei
Publicado: (2024)
por: Liu, Jingwei
Publicado: (2024)
End-to-end Piano Performance-MIDI to Score Conversion with Transformers
por: Beyer, Tim, et al.
Publicado: (2024)
por: Beyer, Tim, et al.
Publicado: (2024)
Notochord: a Flexible Probabilistic Model for Real-Time MIDI Performance
por: Shepardson, Victor, et al.
Publicado: (2024)
por: Shepardson, Victor, et al.
Publicado: (2024)
SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
por: Huang, Wen-Chin, et al.
Publicado: (2025)
por: Huang, Wen-Chin, et al.
Publicado: (2025)
MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
por: Huang, Wen-Chin, et al.
Publicado: (2024)
por: Huang, Wen-Chin, et al.
Publicado: (2024)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
por: Huang, Wen-Chin, et al.
Publicado: (2026)
por: Huang, Wen-Chin, et al.
Publicado: (2026)
Moises-Light: Resource-efficient Band-split U-Net For Music Source Separation
por: Yun-Ning, et al.
Publicado: (2025)
por: Yun-Ning, et al.
Publicado: (2025)
Beat-Based Rhythm Quantization of MIDI Performances
por: Wachter, Maximilian, et al.
Publicado: (2025)
por: Wachter, Maximilian, et al.
Publicado: (2025)
Annotation-Free MIDI-to-Audio Synthesis via Concatenative Synthesis and Generative Refinement
por: Take, Osamu, et al.
Publicado: (2024)
por: Take, Osamu, et al.
Publicado: (2024)
BreathNet: Generalizable Audio Deepfake Detection via Breath-Cue-Guided Feature Refinement
por: Ye, Zhe, et al.
Publicado: (2026)
por: Ye, Zhe, et al.
Publicado: (2026)
Reproducing the Acoustic Velocity Vectors in a Circular Listening Area
por: Wang, Jiarui, et al.
Publicado: (2024)
por: Wang, Jiarui, et al.
Publicado: (2024)
Sub-band and Full-band Interactive U-Net with DPRNN for Demixing Cross-talk Stereo Music
por: Yin, Han, et al.
Publicado: (2024)
por: Yin, Han, et al.
Publicado: (2024)
An Adaptive CMSA for Solving the Longest Filled Common Subsequence Problem with an Application in Audio Querying
por: Djukanovic, Marko, et al.
Publicado: (2025)
por: Djukanovic, Marko, et al.
Publicado: (2025)
Moonbeam: A MIDI Foundation Model Using Both Absolute and Relative Music Attributes
por: Guo, Zixun, et al.
Publicado: (2025)
por: Guo, Zixun, et al.
Publicado: (2025)
COVID-19 Diagnosis from Cough Acoustics using ConvNets and Data Augmentation
por: Mahanta, Saranga Kingkor, et al.
Publicado: (2021)
por: Mahanta, Saranga Kingkor, et al.
Publicado: (2021)
MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling
por: Tang, Jingjing, et al.
Publicado: (2025)
por: Tang, Jingjing, et al.
Publicado: (2025)
Composer's Assistant 2: Interactive Multi-Track MIDI Infilling with Fine-Grained User Control
por: Malandro, Martin E.
Publicado: (2024)
por: Malandro, Martin E.
Publicado: (2024)
Fretting-Transformer: Encoder-Decoder Model for MIDI to Tablature Transcription
por: Hamberger, Anna, et al.
Publicado: (2025)
por: Hamberger, Anna, et al.
Publicado: (2025)
SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
por: Gu, Yicheng, et al.
Publicado: (2025)
por: Gu, Yicheng, et al.
Publicado: (2025)
AuralNet: Hierarchical Attention-based 3D Binaural Localization of Overlapping Speakers
por: Fu, Linya, et al.
Publicado: (2025)
por: Fu, Linya, et al.
Publicado: (2025)
Fine-Tuning MIDI-to-Audio Alignment using a Neural Network on Piano Roll and CQT Representations
por: Murgul, Sebastian, et al.
Publicado: (2025)
por: Murgul, Sebastian, et al.
Publicado: (2025)
RaD-Net: A Repairing and Denoising Network for Speech Signal Improvement
por: Liu, Mingshuai, et al.
Publicado: (2024)
por: Liu, Mingshuai, et al.
Publicado: (2024)
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
por: Jeong, Jaeseok, et al.
Publicado: (2025)
por: Jeong, Jaeseok, et al.
Publicado: (2025)
GMM-ResNet2: Ensemble of Group ResNet Networks for Synthetic Speech Detection
por: Lei, Zhenchun, et al.
Publicado: (2024)
por: Lei, Zhenchun, et al.
Publicado: (2024)
MidiCaps: A large-scale MIDI dataset with text captions
por: Melechovsky, Jan, et al.
Publicado: (2024)
por: Melechovsky, Jan, et al.
Publicado: (2024)
PhiNet: Speaker Verification with Phonetic Interpretability
por: Ma, Yi, et al.
Publicado: (2026)
por: Ma, Yi, et al.
Publicado: (2026)
KS-Net: Multi-band joint speech restoration and enhancement network for 2024 ICASSP SSI Challenge
por: Yu, Guochen, et al.
Publicado: (2024)
por: Yu, Guochen, et al.
Publicado: (2024)
Beat and Downbeat Tracking in Performance MIDI Using an End-to-End Transformer Architecture
por: Murgul, Sebastian, et al.
Publicado: (2025)
por: Murgul, Sebastian, et al.
Publicado: (2025)
Asynchronous Microphone Array Calibration using Hybrid TDOA Information
por: Zhang, Chengjie, et al.
Publicado: (2024)
por: Zhang, Chengjie, et al.
Publicado: (2024)
PlumberNet: Fixing interference leakage after GEV beamforming
por: Grondin, François, et al.
Publicado: (2023)
por: Grondin, François, et al.
Publicado: (2023)
Inter-channel Conv-TasNet for multichannel speech enhancement
por: Lee, Dongheon, et al.
Publicado: (2021)
por: Lee, Dongheon, et al.
Publicado: (2021)
Ejemplares similares
-
Score-Informed Transformer for Refining MIDI Velocity in Automatic Music Transcription
por: He, Zhanhong, et al.
Publicado: (2025) -
How to Infer Repeat Structures in MIDI Performances
por: Peter, Silvan, et al.
Publicado: (2025) -
Pseudo Strong Labels from Frame-Level Predictions for Weakly Supervised Sound Event Detection
por: Zhang, Yuliang, et al.
Publicado: (2025) -
Impact of Noisy Labels on Sound Event Detection: Deletion Errors Are More Detrimental Than Insertion Errors
por: Zhang, Yuliang, et al.
Publicado: (2024) -
MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline
por: Tsai, Fang-Duo, et al.
Publicado: (2026)