MidiCaps: A large-scale MIDI dataset with text captions
Fuente:
arXiv
Salvato in:
| Autori principali: | Melechovsky, Jan, Roy, Abhinaba, Herremans, Dorien |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
di: Liu, Renhang, et al.
Pubblicazione: (2024)
di: Liu, Renhang, et al.
Pubblicazione: (2024)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
di: Roy, Abhinaba, et al.
Pubblicazione: (2025)
di: Roy, Abhinaba, et al.
Pubblicazione: (2025)
On the de-duplication of the Lakh MIDI dataset
di: Choi, Eunjin, et al.
Pubblicazione: (2025)
di: Choi, Eunjin, et al.
Pubblicazione: (2025)
MelodySim: Measuring Melody-aware Music Similarity for Plagiarism Detection
di: Lu, Tongyu, et al.
Pubblicazione: (2025)
di: Lu, Tongyu, et al.
Pubblicazione: (2025)
Aligning Generative Music AI with Human Preferences: Methods and Challenges
di: Herremans, Dorien, et al.
Pubblicazione: (2025)
di: Herremans, Dorien, et al.
Pubblicazione: (2025)
MidiTok Visualizer: a tool for visualization and analysis of tokenized MIDI symbolic music
di: Wiszenko, Michał, et al.
Pubblicazione: (2024)
di: Wiszenko, Michał, et al.
Pubblicazione: (2024)
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
di: Melechovsky, Jan, et al.
Pubblicazione: (2025)
di: Melechovsky, Jan, et al.
Pubblicazione: (2025)
MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music Composition
di: Pasquier, Philippe, et al.
Pubblicazione: (2025)
di: Pasquier, Philippe, et al.
Pubblicazione: (2025)
BandCondiNet: Parallel Transformers-based Conditional Popular Music Generation with Multi-View Features
di: Luo, Jing, et al.
Pubblicazione: (2024)
di: Luo, Jing, et al.
Pubblicazione: (2024)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
Dance2MIDI: Dance-driven multi-instruments music generation
di: Han, Bo, et al.
Pubblicazione: (2023)
di: Han, Bo, et al.
Pubblicazione: (2023)
Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations
di: Wachter, Maximilian, et al.
Pubblicazione: (2026)
di: Wachter, Maximilian, et al.
Pubblicazione: (2026)
Beat-Based Rhythm Quantization of MIDI Performances
di: Wachter, Maximilian, et al.
Pubblicazione: (2025)
di: Wachter, Maximilian, et al.
Pubblicazione: (2025)
A Traditional Approach to Symbolic Piano Continuation
di: Zhou-Zheng, Christian, et al.
Pubblicazione: (2025)
di: Zhou-Zheng, Christian, et al.
Pubblicazione: (2025)
CHORDONOMICON: A Dataset of 666,000 Songs and their Chord Progressions
di: Kantarelis, Spyridon, et al.
Pubblicazione: (2024)
di: Kantarelis, Spyridon, et al.
Pubblicazione: (2024)
A Recurrent Neural Network Approach to the Answering Machine Detection Problem
di: Altwlkany, Kemal, et al.
Pubblicazione: (2024)
di: Altwlkany, Kemal, et al.
Pubblicazione: (2024)
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
di: Sadok, Samir, et al.
Pubblicazione: (2023)
di: Sadok, Samir, et al.
Pubblicazione: (2023)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
di: Sadok, Samir, et al.
Pubblicazione: (2023)
di: Sadok, Samir, et al.
Pubblicazione: (2023)
A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
di: Ishii, Masato, et al.
Pubblicazione: (2024)
di: Ishii, Masato, et al.
Pubblicazione: (2024)
Fretting-Transformer: Encoder-Decoder Model for MIDI to Tablature Transcription
di: Hamberger, Anna, et al.
Pubblicazione: (2025)
di: Hamberger, Anna, et al.
Pubblicazione: (2025)
TEAdapter: Supply abundant guidance for controllable text-to-music generation
di: Zou, Jialing, et al.
Pubblicazione: (2024)
di: Zou, Jialing, et al.
Pubblicazione: (2024)
Exploring compressibility of transformer based text-to-music (TTM) models
di: Moschopoulos, Vasileios, et al.
Pubblicazione: (2024)
di: Moschopoulos, Vasileios, et al.
Pubblicazione: (2024)
Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
di: Bandyopadhyay, Tathagata
Pubblicazione: (2024)
di: Bandyopadhyay, Tathagata
Pubblicazione: (2024)
Siamese Residual Neural Network for Musical Shape Evaluation in Piano Performance Assessment
di: Li, Xiaoquan, et al.
Pubblicazione: (2024)
di: Li, Xiaoquan, et al.
Pubblicazione: (2024)
Efficient Feature Extraction and Late Fusion Strategy for Audiovisual Emotional Mimicry Intensity Estimation
di: Yu, Jun, et al.
Pubblicazione: (2024)
di: Yu, Jun, et al.
Pubblicazione: (2024)
Compression of Higher Order Ambisonics with Multichannel RVQGAN
di: Hirvonen, Toni, et al.
Pubblicazione: (2024)
di: Hirvonen, Toni, et al.
Pubblicazione: (2024)
Source Separation of Multi-source Raw Music using a Residual Quantized Variational Autoencoder
di: Berti, Leonardo
Pubblicazione: (2024)
di: Berti, Leonardo
Pubblicazione: (2024)
Audiopedia: Audio QA with Knowledge
di: Penamakuri, Abhirama Subramanyam, et al.
Pubblicazione: (2024)
di: Penamakuri, Abhirama Subramanyam, et al.
Pubblicazione: (2024)
LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement
di: Jain, Arnav, et al.
Pubblicazione: (2024)
di: Jain, Arnav, et al.
Pubblicazione: (2024)
Unified Microphone Conversion: Many-to-Many Device Mapping via Feature-wise Linear Modulation
di: Ryu, Myeonghoon, et al.
Pubblicazione: (2024)
di: Ryu, Myeonghoon, et al.
Pubblicazione: (2024)
Microphone Conversion: Mitigating Device Variability in Sound Event Classification
di: Ryu, Myeonghoon, et al.
Pubblicazione: (2024)
di: Ryu, Myeonghoon, et al.
Pubblicazione: (2024)
Just Label the Repeats for In-The-Wild Audio-to-Score Alignment
di: Bukey, Irmak, et al.
Pubblicazione: (2024)
di: Bukey, Irmak, et al.
Pubblicazione: (2024)
Music102: An $D_{12}$-equivariant transformer for chord progression accompaniment
di: Luo, Weiliang
Pubblicazione: (2024)
di: Luo, Weiliang
Pubblicazione: (2024)
Network Bending of Diffusion Models for Audio-Visual Generation
di: Dzwonczyk, Luke, et al.
Pubblicazione: (2024)
di: Dzwonczyk, Luke, et al.
Pubblicazione: (2024)
Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
di: Wang, Wupeng, et al.
Pubblicazione: (2024)
di: Wang, Wupeng, et al.
Pubblicazione: (2024)
AWARE: Audio Watermarking with Adversarial Resistance to Edits
di: Pavlović, Kosta, et al.
Pubblicazione: (2025)
di: Pavlović, Kosta, et al.
Pubblicazione: (2025)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
di: Lin, Meng-Ping, et al.
Pubblicazione: (2025)
di: Lin, Meng-Ping, et al.
Pubblicazione: (2025)
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
di: Herbuela, Von Ralph Dane Marquez, et al.
Pubblicazione: (2025)
di: Herbuela, Von Ralph Dane Marquez, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
di: Liu, Renhang, et al.
Pubblicazione: (2024) -
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
di: Melechovsky, Jan, et al.
Pubblicazione: (2024) -
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
di: Melechovsky, Jan, et al.
Pubblicazione: (2022) -
Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
di: Roy, Abhinaba, et al.
Pubblicazione: (2025) -
On the de-duplication of the Lakh MIDI dataset
di: Choi, Eunjin, et al.
Pubblicazione: (2025)