LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Chou, Benjamin Shiue-Hal, Jajal, Purvish, Eliopoulos, Nick John, Davis, James C., Thiruvathukal, George K., Yun, Kristen Yeon-Ji, Lu, Yung-Hsiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting Music Performance Errors with Transformers
by: Chou, Benjamin Shiue-Hal, et al.
Published: (2025)
by: Chou, Benjamin Shiue-Hal, et al.
Published: (2025)
Inference-Time Alignment of Diffusion Models via Evolutionary Algorithms
by: Jajal, Purvish, et al.
Published: (2025)
by: Jajal, Purvish, et al.
Published: (2025)
Token Turing Machines are Efficient Vision Models
by: Jajal, Purvish, et al.
Published: (2024)
by: Jajal, Purvish, et al.
Published: (2024)
AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
by: Jajal, Purvish, et al.
Published: (2025)
by: Jajal, Purvish, et al.
Published: (2025)
SymPAC: Scalable Symbolic Music Generation With Prompts And Constraints
by: Chen, Haonan, et al.
Published: (2024)
by: Chen, Haonan, et al.
Published: (2024)
Towards Blind Data Cleaning: A Case Study in Music Source Separation
by: Gui, Azalea, et al.
Published: (2025)
by: Gui, Azalea, et al.
Published: (2025)
The Rarity of Musical Audio Signals Within the Space of Possible Audio Generation
by: Collins, Nick
Published: (2024)
by: Collins, Nick
Published: (2024)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
by: Niu, Xinlei, et al.
Published: (2025)
by: Niu, Xinlei, et al.
Published: (2025)
ITO-Master: Inference-Time Optimization for Audio Effects Modeling of Music Mastering Processors
by: Koo, Junghyun, et al.
Published: (2025)
by: Koo, Junghyun, et al.
Published: (2025)
WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation
by: Karystinaios, Emmanouil
Published: (2025)
by: Karystinaios, Emmanouil
Published: (2025)
A Survey on Multimodal Music Emotion Recognition
by: Liyanarachchi, Rashini, et al.
Published: (2025)
by: Liyanarachchi, Rashini, et al.
Published: (2025)
Automatic Music Mixing using a Generative Model of Effect Embeddings
by: Moliner, Eloi, et al.
Published: (2025)
by: Moliner, Eloi, et al.
Published: (2025)
Reducing Barriers to the Use of Marginalised Music Genres in AI
by: Bryan-Kinns, Nick, et al.
Published: (2024)
by: Bryan-Kinns, Nick, et al.
Published: (2024)
Pre-training Music Classification Models via Music Source Separation
by: Garoufis, Christos, et al.
Published: (2023)
by: Garoufis, Christos, et al.
Published: (2023)
Moises-Light: Resource-efficient Band-split U-Net For Music Source Separation
by: Yun-Ning, et al.
Published: (2025)
by: Yun-Ning, et al.
Published: (2025)
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
by: Choi, Woosung, et al.
Published: (2025)
by: Choi, Woosung, et al.
Published: (2025)
AI in Music and Sound: Pedagogical Reflections, Post-Structuralist Approaches and Creative Outcomes in Seminar Practice
by: Coelho, Guilherme
Published: (2025)
by: Coelho, Guilherme
Published: (2025)
Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
by: Tang, Jingjing, et al.
Published: (2025)
by: Tang, Jingjing, et al.
Published: (2025)
ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism
by: Chou, Hsing-Hang, et al.
Published: (2024)
by: Chou, Hsing-Hang, et al.
Published: (2024)
Musical Word Embedding for Music Tagging and Retrieval
by: Doh, SeungHeon, et al.
Published: (2024)
by: Doh, SeungHeon, et al.
Published: (2024)
Expressive Timing in Hindustani Vocal Music
by: Bhake, Yash, et al.
Published: (2025)
by: Bhake, Yash, et al.
Published: (2025)
SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
by: Yang, Chenyu, et al.
Published: (2025)
by: Yang, Chenyu, et al.
Published: (2025)
Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
by: Batlle-Roca, Roser, et al.
Published: (2024)
by: Batlle-Roca, Roser, et al.
Published: (2024)
AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances
by: Zurale, Devansh, et al.
Published: (2026)
by: Zurale, Devansh, et al.
Published: (2026)
MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
by: Liu, Cheng, et al.
Published: (2025)
by: Liu, Cheng, et al.
Published: (2025)
A Survey on Cross-Modal Interaction Between Music and Multimodal Data
by: Li, Sifei, et al.
Published: (2025)
by: Li, Sifei, et al.
Published: (2025)
MusFlow: Multimodal Music Generation via Conditional Flow Matching
by: Song, Jiahao, et al.
Published: (2025)
by: Song, Jiahao, et al.
Published: (2025)
Semi-Supervised Contrastive Learning of Musical Representations
by: Guinot, Julien, et al.
Published: (2024)
by: Guinot, Julien, et al.
Published: (2024)
Mustango: Toward Controllable Text-to-Music Generation
by: Melechovsky, Jan, et al.
Published: (2023)
by: Melechovsky, Jan, et al.
Published: (2023)
Music2Fail: Transfer Music to Failed Recorder Style
by: Leong, Chon In, et al.
Published: (2024)
by: Leong, Chon In, et al.
Published: (2024)
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
by: Comanducci, Luca, et al.
Published: (2024)
by: Comanducci, Luca, et al.
Published: (2024)
Revisiting Modeling and Evaluation Approaches in Speech Emotion Recognition: Considering Subjectivity of Annotators and Ambiguity of Emotions
by: Chou, Huang-Cheng, et al.
Published: (2025)
by: Chou, Huang-Cheng, et al.
Published: (2025)
SCNet: Sparse Compression Network for Music Source Separation
by: Tong, Weinan, et al.
Published: (2024)
by: Tong, Weinan, et al.
Published: (2024)
Multi-Distillation from Speech and Music Representation Models
by: Wei, Jui-Chiang, et al.
Published: (2025)
by: Wei, Jui-Chiang, et al.
Published: (2025)
Identification and Clustering of Unseen Ragas in Indian Art Music
by: Singh, Parampreet, et al.
Published: (2024)
by: Singh, Parampreet, et al.
Published: (2024)
Adapting Frechet Audio Distance for Generative Music Evaluation
by: Gui, Azalea, et al.
Published: (2023)
by: Gui, Azalea, et al.
Published: (2023)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
by: Weck, Benno, et al.
Published: (2024)
by: Weck, Benno, et al.
Published: (2024)
Semantic-Aware Interpretable Multimodal Music Auto-Tagging
by: Patakis, Andreas, et al.
Published: (2025)
by: Patakis, Andreas, et al.
Published: (2025)
TALKPLAY: Multimodal Music Recommendation with Large Language Models
by: Doh, Seungheon, et al.
Published: (2025)
by: Doh, Seungheon, et al.
Published: (2025)
Similar Items
-
Detecting Music Performance Errors with Transformers
by: Chou, Benjamin Shiue-Hal, et al.
Published: (2025) -
Inference-Time Alignment of Diffusion Models via Evolutionary Algorithms
by: Jajal, Purvish, et al.
Published: (2025) -
Token Turing Machines are Efficient Vision Models
by: Jajal, Purvish, et al.
Published: (2024) -
AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
by: Jajal, Purvish, et al.
Published: (2025) -
SymPAC: Scalable Symbolic Music Generation With Prompts And Constraints
by: Chen, Haonan, et al.
Published: (2024)