From Image to Music Language: A Two-Stage Structure Decoding Approach for Complex Polyphonic OMR
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Nan, Li, Shiheng, Hou, Shengchao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Discovering "Words" in Music: Unsupervised Learning of Compositional Sparse Code for Symbolic Music
por: Wang, Tianle, et al.
Publicado: (2025)
por: Wang, Tianle, et al.
Publicado: (2025)
TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation
por: Liu, Xinran, et al.
Publicado: (2026)
por: Liu, Xinran, et al.
Publicado: (2026)
Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation
por: Lyu, Guangtao, et al.
Publicado: (2025)
por: Lyu, Guangtao, et al.
Publicado: (2025)
Music Genre Classification using Large Language Models
por: Meguenani, Mohamed El Amine, et al.
Publicado: (2024)
por: Meguenani, Mohamed El Amine, et al.
Publicado: (2024)
MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
por: Gupta, Prerit, et al.
Publicado: (2025)
por: Gupta, Prerit, et al.
Publicado: (2025)
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
por: Li, You, et al.
Publicado: (2026)
por: Li, You, et al.
Publicado: (2026)
SonoWorld: From One Image to a 3D Audio-Visual Scene
por: Jin, Derong, et al.
Publicado: (2026)
por: Jin, Derong, et al.
Publicado: (2026)
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
por: Liu, Zehua, et al.
Publicado: (2025)
por: Liu, Zehua, et al.
Publicado: (2025)
MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
por: Chi, Xiaowei, et al.
Publicado: (2024)
por: Chi, Xiaowei, et al.
Publicado: (2024)
DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
por: Zhang, Hengyuan, et al.
Publicado: (2025)
por: Zhang, Hengyuan, et al.
Publicado: (2025)
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
por: Rinaldi, Ivan, et al.
Publicado: (2026)
por: Rinaldi, Ivan, et al.
Publicado: (2026)
FLUX that Plays Music
por: Fei, Zhengcong, et al.
Publicado: (2024)
por: Fei, Zhengcong, et al.
Publicado: (2024)
Exploring Multi-Modal Control in Music-Driven Dance Generation
por: Li, Ronghui, et al.
Publicado: (2024)
por: Li, Ronghui, et al.
Publicado: (2024)
Emotion-Guided Image to Music Generation
por: Kundu, Souraja, et al.
Publicado: (2024)
por: Kundu, Souraja, et al.
Publicado: (2024)
Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription
por: Ríos-Vila, Antonio, et al.
Publicado: (2024)
por: Ríos-Vila, Antonio, et al.
Publicado: (2024)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
por: Kim, Inho, et al.
Publicado: (2025)
por: Kim, Inho, et al.
Publicado: (2025)
Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment
por: Hong, Jiaying, et al.
Publicado: (2025)
por: Hong, Jiaying, et al.
Publicado: (2025)
TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba
por: Yang, Ziyue, et al.
Publicado: (2026)
por: Yang, Ziyue, et al.
Publicado: (2026)
DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling
por: Ghosh, Anindita, et al.
Publicado: (2025)
por: Ghosh, Anindita, et al.
Publicado: (2025)
Knowledge Discovery in Optical Music Recognition: Enhancing Information Retrieval with Instance Segmentation
por: Shatri, Elona, et al.
Publicado: (2024)
por: Shatri, Elona, et al.
Publicado: (2024)
VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
por: Tian, Zeyue, et al.
Publicado: (2024)
por: Tian, Zeyue, et al.
Publicado: (2024)
DanceChat: Large Language Model-Guided Music-to-Dance Generation
por: Wang, Qing, et al.
Publicado: (2025)
por: Wang, Qing, et al.
Publicado: (2025)
Learning Musical Representations for Music Performance Question Answering
por: Diao, Xingjian, et al.
Publicado: (2025)
por: Diao, Xingjian, et al.
Publicado: (2025)
Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
por: Chuchra, Akanksha, et al.
Publicado: (2026)
por: Chuchra, Akanksha, et al.
Publicado: (2026)
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
por: Su, Yaofeng, et al.
Publicado: (2026)
por: Su, Yaofeng, et al.
Publicado: (2026)
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
por: Zhou, Yupeng, et al.
Publicado: (2026)
por: Zhou, Yupeng, et al.
Publicado: (2026)
VABench: A Comprehensive Benchmark for Audio-Video Generation
por: Hua, Daili, et al.
Publicado: (2025)
por: Hua, Daili, et al.
Publicado: (2025)
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
por: Zhang, Zeliang, et al.
Publicado: (2025)
por: Zhang, Zeliang, et al.
Publicado: (2025)
VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
por: Lin, Yan-Bo, et al.
Publicado: (2024)
por: Lin, Yan-Bo, et al.
Publicado: (2024)
Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings
por: Hisariya, Tanisha, et al.
Publicado: (2024)
por: Hisariya, Tanisha, et al.
Publicado: (2024)
High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
por: Huang, Chao, et al.
Publicado: (2025)
por: Huang, Chao, et al.
Publicado: (2025)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
por: Tang, Changli, et al.
Publicado: (2025)
por: Tang, Changli, et al.
Publicado: (2025)
Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings
por: Gohari, Mahyar, et al.
Publicado: (2024)
por: Gohari, Mahyar, et al.
Publicado: (2024)
Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
por: Liu, Chen, et al.
Publicado: (2025)
por: Liu, Chen, et al.
Publicado: (2025)
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
por: Li, Kai, et al.
Publicado: (2025)
por: Li, Kai, et al.
Publicado: (2025)
MOVA: Towards Scalable and Synchronized Video-Audio Generation
por: OpenMOSS Team, et al.
Publicado: (2026)
por: OpenMOSS Team, et al.
Publicado: (2026)
Decoding Emotions: Unveiling Facial Expressions through Acoustic Sensing with Contrastive Attention
por: Wang, Guangjing, et al.
Publicado: (2024)
por: Wang, Guangjing, et al.
Publicado: (2024)
Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks
por: Liu, Chang, et al.
Publicado: (2025)
por: Liu, Chang, et al.
Publicado: (2025)
REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation
por: Wang, Haotian, et al.
Publicado: (2025)
por: Wang, Haotian, et al.
Publicado: (2025)
An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits
por: Li, Kai, et al.
Publicado: (2022)
por: Li, Kai, et al.
Publicado: (2022)
Ejemplares similares
-
Discovering "Words" in Music: Unsupervised Learning of Compositional Sparse Code for Symbolic Music
por: Wang, Tianle, et al.
Publicado: (2025) -
TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation
por: Liu, Xinran, et al.
Publicado: (2026) -
Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation
por: Lyu, Guangtao, et al.
Publicado: (2025) -
Music Genre Classification using Large Language Models
por: Meguenani, Mohamed El Amine, et al.
Publicado: (2024) -
MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
por: Gupta, Prerit, et al.
Publicado: (2025)