Harmonizing Pixels and Melodies: Maestro-Guided Film Score Generation and Composition Style Transfer
Fuente:
arXiv
Salvato in:
| Autori principali: | Qi, F., Ni, L., Xu, C. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial Animation
di: Yang, An, et al.
Pubblicazione: (2025)
di: Yang, An, et al.
Pubblicazione: (2025)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
di: Hsiao, Teng-Fang, et al.
Pubblicazione: (2024)
di: Hsiao, Teng-Fang, et al.
Pubblicazione: (2024)
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
di: Wang, Song, et al.
Pubblicazione: (2025)
di: Wang, Song, et al.
Pubblicazione: (2025)
DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
di: Hu, Ying, et al.
Pubblicazione: (2024)
di: Hu, Ying, et al.
Pubblicazione: (2024)
Neural Style Transfer for Audio Spectograms
di: Verma, Prateek, et al.
Pubblicazione: (2018)
di: Verma, Prateek, et al.
Pubblicazione: (2018)
PixelatedScatter: Arbitrary-level Visual Abstraction for Large-scale Multiclass Scatterplots
di: Guo, Ziheng, et al.
Pubblicazione: (2025)
di: Guo, Ziheng, et al.
Pubblicazione: (2025)
MusicWeaver: Composer-Style Structural Editing and Minute-Scale Coherent Music Generation
di: Wang, Xuanchen, et al.
Pubblicazione: (2025)
di: Wang, Xuanchen, et al.
Pubblicazione: (2025)
Inter-Frame Coding for Dynamic Meshes via Coarse-to-Fine Anchor Mesh Generation
di: Huang, He, et al.
Pubblicazione: (2024)
di: Huang, He, et al.
Pubblicazione: (2024)
Harnessing the Latent Diffusion Model for Training-Free Image Style Transfer
di: Masui, Kento, et al.
Pubblicazione: (2024)
di: Masui, Kento, et al.
Pubblicazione: (2024)
Controllable Dance Generation with Style-Guided Motion Diffusion
di: Wang, Hongsong, et al.
Pubblicazione: (2024)
di: Wang, Hongsong, et al.
Pubblicazione: (2024)
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
di: Wang, Yongqi, et al.
Pubblicazione: (2025)
di: Wang, Yongqi, et al.
Pubblicazione: (2025)
StePO-Rec: Towards Personalized Outfit Styling Assistant via Knowledge-Guided Multi-Step Reasoning
di: Bi, Yuxi, et al.
Pubblicazione: (2025)
di: Bi, Yuxi, et al.
Pubblicazione: (2025)
Detecting Notational Errors in Digital Music Scores
di: Léo, Géré, et al.
Pubblicazione: (2025)
di: Léo, Géré, et al.
Pubblicazione: (2025)
Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
di: Zeng, Wei, et al.
Pubblicazione: (2025)
di: Zeng, Wei, et al.
Pubblicazione: (2025)
Latent Feature-Guided Conditional Diffusion for Generative Image Semantic Communication
di: Chen, Zehao, et al.
Pubblicazione: (2025)
di: Chen, Zehao, et al.
Pubblicazione: (2025)
Harmonizing Attention: Training-free Texture-aware Geometry Transfer
di: Ikuta, Eito, et al.
Pubblicazione: (2024)
di: Ikuta, Eito, et al.
Pubblicazione: (2024)
Contribution-Guided Asymmetric Learning for Robust Multimodal Fusion under Imbalance and Noise
di: Xu, Zijing, et al.
Pubblicazione: (2025)
di: Xu, Zijing, et al.
Pubblicazione: (2025)
Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information
di: Huang, Qiaochu, et al.
Pubblicazione: (2024)
di: Huang, Qiaochu, et al.
Pubblicazione: (2024)
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
di: Zi, Xing, et al.
Pubblicazione: (2025)
di: Zi, Xing, et al.
Pubblicazione: (2025)
Multi-scale Attention Guided Pose Transfer
di: Roy, Prasun, et al.
Pubblicazione: (2022)
di: Roy, Prasun, et al.
Pubblicazione: (2022)
Enhancing Film Grain Coding in VVC: Improving Encoding Quality and Efficiency
di: Menon, Vignesh V, et al.
Pubblicazione: (2024)
di: Menon, Vignesh V, et al.
Pubblicazione: (2024)
Protégé: Learn and Generate Basic Makeup Styles with Generative Adversarial Networks (GANs)
di: Sii, Jia Wei, et al.
Pubblicazione: (2024)
di: Sii, Jia Wei, et al.
Pubblicazione: (2024)
MusicScore: A Dataset for Music Score Modeling and Generation
di: Lin, Yuheng, et al.
Pubblicazione: (2024)
di: Lin, Yuheng, et al.
Pubblicazione: (2024)
Gain of Grain: A Film Grain Handling Toolchain for VVC-based Open Implementations
di: Menon, Vignesh V, et al.
Pubblicazione: (2024)
di: Menon, Vignesh V, et al.
Pubblicazione: (2024)
LLM2Manim: Pedagogy-Aware AI Generation of STEM Animations
di: Joshi, Aastha, et al.
Pubblicazione: (2026)
di: Joshi, Aastha, et al.
Pubblicazione: (2026)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025)
di: Wu, Yi, et al.
Pubblicazione: (2025)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
di: Zhang, Juan, et al.
Pubblicazione: (2024)
di: Zhang, Juan, et al.
Pubblicazione: (2024)
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
di: Wang, Xuanchen, et al.
Pubblicazione: (2025)
di: Wang, Xuanchen, et al.
Pubblicazione: (2025)
Generative AI-enabled Mobile Tactical Multimedia Networks: Distribution, Generation, and Perception
di: Xu, Minrui, et al.
Pubblicazione: (2024)
di: Xu, Minrui, et al.
Pubblicazione: (2024)
MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
di: Yuan, Xiang, et al.
Pubblicazione: (2026)
di: Yuan, Xiang, et al.
Pubblicazione: (2026)
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
di: Cai, Qi, et al.
Pubblicazione: (2026)
di: Cai, Qi, et al.
Pubblicazione: (2026)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
di: Huang, Yiheng, et al.
Pubblicazione: (2025)
di: Huang, Yiheng, et al.
Pubblicazione: (2025)
An Inverse Partial Optimal Transport Framework for Music-guided Movie Trailer Generation
di: Wang, Yutong, et al.
Pubblicazione: (2024)
di: Wang, Yutong, et al.
Pubblicazione: (2024)
Design of a 5G Multimedia Broadcast Application Function Supporting Adaptive Error Recovery
di: Lentisco, C. M., et al.
Pubblicazione: (2024)
di: Lentisco, C. M., et al.
Pubblicazione: (2024)
Reducing Latency for Multimedia Broadcast Services Over Mobile Networks
di: Lentisco, C. M., et al.
Pubblicazione: (2024)
di: Lentisco, C. M., et al.
Pubblicazione: (2024)
Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-training Models with Contrastive Learning
di: Wang, Youze, et al.
Pubblicazione: (2023)
di: Wang, Youze, et al.
Pubblicazione: (2023)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
di: Guan, Jiazhi, et al.
Pubblicazione: (2024)
di: Guan, Jiazhi, et al.
Pubblicazione: (2024)
CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential Recommendation
di: Xu, Jinfeng, et al.
Pubblicazione: (2026)
di: Xu, Jinfeng, et al.
Pubblicazione: (2026)
Diffusion Model-Based Size Variable Virtual Try-On Technology and Evaluation Method
di: Zhang, Shufang, et al.
Pubblicazione: (2025)
di: Zhang, Shufang, et al.
Pubblicazione: (2025)
HarmonyIQA: Pioneering Benchmark and Model for Image Harmonization Quality Assessment
di: Xu, Zitong, et al.
Pubblicazione: (2025)
di: Xu, Zitong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial Animation
di: Yang, An, et al.
Pubblicazione: (2025) -
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
di: Hsiao, Teng-Fang, et al.
Pubblicazione: (2024) -
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
di: Wang, Song, et al.
Pubblicazione: (2025) -
DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
di: Hu, Ying, et al.
Pubblicazione: (2024) -
Neural Style Transfer for Audio Spectograms
di: Verma, Prateek, et al.
Pubblicazione: (2018)