Similar Items
StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial Animation
by: Yang, An, et al.
Published: (2025)
by: Yang, An, et al.
Published: (2025)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
by: Hsiao, Teng-Fang, et al.
Published: (2024)
by: Hsiao, Teng-Fang, et al.
Published: (2024)
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
by: Hu, Ying, et al.
Published: (2024)
by: Hu, Ying, et al.
Published: (2024)
Neural Style Transfer for Audio Spectograms
by: Verma, Prateek, et al.
Published: (2018)
by: Verma, Prateek, et al.
Published: (2018)
PixelatedScatter: Arbitrary-level Visual Abstraction for Large-scale Multiclass Scatterplots
by: Guo, Ziheng, et al.
Published: (2025)
by: Guo, Ziheng, et al.
Published: (2025)
MusicWeaver: Composer-Style Structural Editing and Minute-Scale Coherent Music Generation
by: Wang, Xuanchen, et al.
Published: (2025)
by: Wang, Xuanchen, et al.
Published: (2025)
Inter-Frame Coding for Dynamic Meshes via Coarse-to-Fine Anchor Mesh Generation
by: Huang, He, et al.
Published: (2024)
by: Huang, He, et al.
Published: (2024)
Harnessing the Latent Diffusion Model for Training-Free Image Style Transfer
by: Masui, Kento, et al.
Published: (2024)
by: Masui, Kento, et al.
Published: (2024)
Controllable Dance Generation with Style-Guided Motion Diffusion
by: Wang, Hongsong, et al.
Published: (2024)
by: Wang, Hongsong, et al.
Published: (2024)
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
by: Wang, Yongqi, et al.
Published: (2025)
by: Wang, Yongqi, et al.
Published: (2025)
StePO-Rec: Towards Personalized Outfit Styling Assistant via Knowledge-Guided Multi-Step Reasoning
by: Bi, Yuxi, et al.
Published: (2025)
by: Bi, Yuxi, et al.
Published: (2025)
Detecting Notational Errors in Digital Music Scores
by: Léo, Géré, et al.
Published: (2025)
by: Léo, Géré, et al.
Published: (2025)
Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
by: Zeng, Wei, et al.
Published: (2025)
by: Zeng, Wei, et al.
Published: (2025)
Latent Feature-Guided Conditional Diffusion for Generative Image Semantic Communication
by: Chen, Zehao, et al.
Published: (2025)
by: Chen, Zehao, et al.
Published: (2025)
Harmonizing Attention: Training-free Texture-aware Geometry Transfer
by: Ikuta, Eito, et al.
Published: (2024)
by: Ikuta, Eito, et al.
Published: (2024)
Contribution-Guided Asymmetric Learning for Robust Multimodal Fusion under Imbalance and Noise
by: Xu, Zijing, et al.
Published: (2025)
by: Xu, Zijing, et al.
Published: (2025)
Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information
by: Huang, Qiaochu, et al.
Published: (2024)
by: Huang, Qiaochu, et al.
Published: (2024)
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
by: Zi, Xing, et al.
Published: (2025)
by: Zi, Xing, et al.
Published: (2025)
Multi-scale Attention Guided Pose Transfer
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
Enhancing Film Grain Coding in VVC: Improving Encoding Quality and Efficiency
by: Menon, Vignesh V, et al.
Published: (2024)
by: Menon, Vignesh V, et al.
Published: (2024)
Protégé: Learn and Generate Basic Makeup Styles with Generative Adversarial Networks (GANs)
by: Sii, Jia Wei, et al.
Published: (2024)
by: Sii, Jia Wei, et al.
Published: (2024)
MusicScore: A Dataset for Music Score Modeling and Generation
by: Lin, Yuheng, et al.
Published: (2024)
by: Lin, Yuheng, et al.
Published: (2024)
Gain of Grain: A Film Grain Handling Toolchain for VVC-based Open Implementations
by: Menon, Vignesh V, et al.
Published: (2024)
by: Menon, Vignesh V, et al.
Published: (2024)
LLM2Manim: Pedagogy-Aware AI Generation of STEM Animations
by: Joshi, Aastha, et al.
Published: (2026)
by: Joshi, Aastha, et al.
Published: (2026)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
by: Wu, Yi, et al.
Published: (2025)
by: Wu, Yi, et al.
Published: (2025)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
by: Zhang, Juan, et al.
Published: (2024)
by: Zhang, Juan, et al.
Published: (2024)
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
by: Wang, Xuanchen, et al.
Published: (2025)
by: Wang, Xuanchen, et al.
Published: (2025)
Generative AI-enabled Mobile Tactical Multimedia Networks: Distribution, Generation, and Perception
by: Xu, Minrui, et al.
Published: (2024)
by: Xu, Minrui, et al.
Published: (2024)
MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
by: Yuan, Xiang, et al.
Published: (2026)
by: Yuan, Xiang, et al.
Published: (2026)
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
by: Cai, Qi, et al.
Published: (2026)
by: Cai, Qi, et al.
Published: (2026)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025)
by: Huang, Yiheng, et al.
Published: (2025)
An Inverse Partial Optimal Transport Framework for Music-guided Movie Trailer Generation
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
Design of a 5G Multimedia Broadcast Application Function Supporting Adaptive Error Recovery
by: Lentisco, C. M., et al.
Published: (2024)
by: Lentisco, C. M., et al.
Published: (2024)
Reducing Latency for Multimedia Broadcast Services Over Mobile Networks
by: Lentisco, C. M., et al.
Published: (2024)
by: Lentisco, C. M., et al.
Published: (2024)
Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-training Models with Contrastive Learning
by: Wang, Youze, et al.
Published: (2023)
by: Wang, Youze, et al.
Published: (2023)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
by: Guan, Jiazhi, et al.
Published: (2024)
by: Guan, Jiazhi, et al.
Published: (2024)
CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential Recommendation
by: Xu, Jinfeng, et al.
Published: (2026)
by: Xu, Jinfeng, et al.
Published: (2026)
Diffusion Model-Based Size Variable Virtual Try-On Technology and Evaluation Method
by: Zhang, Shufang, et al.
Published: (2025)
by: Zhang, Shufang, et al.
Published: (2025)
HarmonyIQA: Pioneering Benchmark and Model for Image Harmonization Quality Assessment
by: Xu, Zitong, et al.
Published: (2025)
by: Xu, Zitong, et al.
Published: (2025)
Similar Items
-
StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial Animation
by: Yang, An, et al.
Published: (2025) -
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
by: Hsiao, Teng-Fang, et al.
Published: (2024) -
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
by: Wang, Song, et al.
Published: (2025) -
DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
by: Hu, Ying, et al.
Published: (2024) -
Neural Style Transfer for Audio Spectograms
by: Verma, Prateek, et al.
Published: (2018)