StemGen: A music generation model that listens
Fuente:
arXiv
Saved in:
| Main Authors: | Parker, Julian D., Spijkervet, Janne, Kosta, Katerina, Yesiler, Furkan, Kuznetsov, Boris, Wang, Ju-Chiang, Avent, Matt, Chen, Jitong, Le, Duc |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
by: Rouard, Simon, et al.
Published: (2025)
by: Rouard, Simon, et al.
Published: (2025)
Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
by: Wang, Ju-Chiang, et al.
Published: (2024)
by: Wang, Ju-Chiang, et al.
Published: (2024)
SymPAC: Scalable Symbolic Music Generation With Prompts And Constraints
by: Chen, Haonan, et al.
Published: (2024)
by: Chen, Haonan, et al.
Published: (2024)
Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis
by: Zhang, Yixiao, et al.
Published: (2025)
by: Zhang, Yixiao, et al.
Published: (2025)
Long-form music generation with latent diffusion
by: Evans, Zach, et al.
Published: (2024)
by: Evans, Zach, et al.
Published: (2024)
Seed-Music: A Unified Framework for High Quality and Controlled Music Generation
by: Bai, Ye, et al.
Published: (2024)
by: Bai, Ye, et al.
Published: (2024)
STASE: A spatialized text-to-audio synthesis engine for music generation
by: Chi, Tutti, et al.
Published: (2025)
by: Chi, Tutti, et al.
Published: (2025)
The first Cadenza challenges: using machine learning competitions to improve music for listeners with a hearing loss
by: Dabike, Gerardo Roa, et al.
Published: (2024)
by: Dabike, Gerardo Roa, et al.
Published: (2024)
Effects of auditory distance cues and reverberation on spatial perception and listening strategies
by: Missoni, Fulvio, et al.
Published: (2025)
by: Missoni, Fulvio, et al.
Published: (2025)
Distilling a speech and music encoder with task arithmetic
by: Ritter-Gutierrez, Fabian, et al.
Published: (2025)
by: Ritter-Gutierrez, Fabian, et al.
Published: (2025)
PiCoGen2: Piano cover generation with transfer learning approach and weakly aligned data
by: Tan, Chih-Pin, et al.
Published: (2024)
by: Tan, Chih-Pin, et al.
Published: (2024)
What do neural networks listen to? Exploring the crucial bands in Speech Enhancement using Sinc-convolution
by: Ho, Kuan-Hsun, et al.
Published: (2024)
by: Ho, Kuan-Hsun, et al.
Published: (2024)
TEAdapter: Supply abundant guidance for controllable text-to-music generation
by: Zou, Jialing, et al.
Published: (2024)
by: Zou, Jialing, et al.
Published: (2024)
Dance2MIDI: Dance-driven multi-instruments music generation
by: Han, Bo, et al.
Published: (2023)
by: Han, Bo, et al.
Published: (2023)
Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
by: Lv, Yishan, et al.
Published: (2026)
by: Lv, Yishan, et al.
Published: (2026)
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
by: Le, Khanh, et al.
Published: (2025)
by: Le, Khanh, et al.
Published: (2025)
Deep learning for music generation. Four approaches and their comparative evaluation
by: Paroiu, Razvan, et al.
Published: (2025)
by: Paroiu, Razvan, et al.
Published: (2025)
Effect of laboratory conditions on the perception of virtual stages for music
by: Accolti, Ernesto
Published: (2025)
by: Accolti, Ernesto
Published: (2025)
Speech foundation models on intelligibility prediction for hearing-impaired listeners
by: Cuervo, Santiago, et al.
Published: (2024)
by: Cuervo, Santiago, et al.
Published: (2024)
STAGE: Stemmed Accompaniment Generation through Prefix-Based Conditioning
by: Strano, Giorgio, et al.
Published: (2025)
by: Strano, Giorgio, et al.
Published: (2025)
LiveScaler: Live control of the harmony of an electronic music track
by: Rixte, Alice
Published: (2024)
by: Rixte, Alice
Published: (2024)
Investigation of perceptual music similarity focusing on each instrumental part
by: Hashizume, Yuka, et al.
Published: (2025)
by: Hashizume, Yuka, et al.
Published: (2025)
Enhanced Automatic Drum Transcription via Drum Stem Source Separation
by: Riley, Xavier, et al.
Published: (2025)
by: Riley, Xavier, et al.
Published: (2025)
PAGURI: a user experience study of creative interaction with text-to-music models
by: Ronchini, Francesca, et al.
Published: (2024)
by: Ronchini, Francesca, et al.
Published: (2024)
PiCoGen: Generate Piano Covers with a Two-stage Approach
by: Tan, Chih-Pin, et al.
Published: (2024)
by: Tan, Chih-Pin, et al.
Published: (2024)
ACMID: Automatic Curation of Musical Instrument Dataset for 7-Stem Music Source Separation
by: Yu, Ji, et al.
Published: (2025)
by: Yu, Ji, et al.
Published: (2025)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
by: Zhao, Junchuan, et al.
Published: (2026)
by: Zhao, Junchuan, et al.
Published: (2026)
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
by: Yao, Jixun, et al.
Published: (2025)
by: Yao, Jixun, et al.
Published: (2025)
SNC: A Stem-Native Codec for Efficient Lossless Audio Storage with Adaptive Playback Capabilities
by: Sufi, Shaad
Published: (2026)
by: Sufi, Shaad
Published: (2026)
Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations
by: Guinot, Julien, et al.
Published: (2024)
by: Guinot, Julien, et al.
Published: (2024)
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
by: Ahmed, Tawsif, et al.
Published: (2025)
by: Ahmed, Tawsif, et al.
Published: (2025)
Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules
by: Chiang, Hsin-Tien, et al.
Published: (2024)
by: Chiang, Hsin-Tien, et al.
Published: (2024)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
by: Le, Khanh, et al.
Published: (2025)
by: Le, Khanh, et al.
Published: (2025)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
by: Le, Khanh, et al.
Published: (2025)
by: Le, Khanh, et al.
Published: (2025)
Room Impulse Responses help attackers to evade Deep Fake Detection
by: Luong, Hieu-Thi, et al.
Published: (2024)
by: Luong, Hieu-Thi, et al.
Published: (2024)
musif: a Python package for symbolic music feature extraction
by: Llorens, Ana, et al.
Published: (2023)
by: Llorens, Ana, et al.
Published: (2023)
Exploring compressibility of transformer based text-to-music (TTM) models
by: Moschopoulos, Vasileios, et al.
Published: (2024)
by: Moschopoulos, Vasileios, et al.
Published: (2024)
Exploring the User Experience of AI-Assisted Sound Searching Systems for Creative Workflows
by: Liu, Haohe, et al.
Published: (2025)
by: Liu, Haohe, et al.
Published: (2025)
Xi+: Uncertainty Supervision for Robust Speaker Embedding
by: Li, Junjie, et al.
Published: (2025)
by: Li, Junjie, et al.
Published: (2025)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
by: Chary, Podakanti Satyajith
Published: (2024)
by: Chary, Podakanti Satyajith
Published: (2024)
Similar Items
-
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
by: Rouard, Simon, et al.
Published: (2025) -
Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
by: Wang, Ju-Chiang, et al.
Published: (2024) -
SymPAC: Scalable Symbolic Music Generation With Prompts And Constraints
by: Chen, Haonan, et al.
Published: (2024) -
Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis
by: Zhang, Yixiao, et al.
Published: (2025) -
Long-form music generation with latent diffusion
by: Evans, Zach, et al.
Published: (2024)