FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Jiahao, Ma, Zhiyong, Du, Wenbiao, Chuai, Qingyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Think Bright, Diffuse Nice: Enhancing T2I-ICL via Inductive-Bias Hint Instruction and Query Contrastive Decoding
by: Ma, Zhiyong, et al.
Published: (2026)
by: Ma, Zhiyong, et al.
Published: (2026)
FlexGen: Flexible Multi-View Generation from Text and Image Inputs
by: Xu, Xinli, et al.
Published: (2024)
by: Xu, Xinli, et al.
Published: (2024)
FSDENet: A Frequency and Spatial Domains based Detail Enhancement Network for Remote Sensing Semantic Segmentation
by: Fu, Jiahao, et al.
Published: (2025)
by: Fu, Jiahao, et al.
Published: (2025)
PiPa++: Towards Unification of Domain Adaptive Semantic Segmentation via Self-supervised Learning
by: Chen, Mu, et al.
Published: (2024)
by: Chen, Mu, et al.
Published: (2024)
Intelligent Parsing: An Automated Parsing Framework for Extracting Design Semantics from E-commerce Creatives
by: Li, Guandong, et al.
Published: (2023)
by: Li, Guandong, et al.
Published: (2023)
Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models
by: Yang, Songlin, et al.
Published: (2026)
by: Yang, Songlin, et al.
Published: (2026)
TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
by: Ma, Zhiyong, et al.
Published: (2025)
by: Ma, Zhiyong, et al.
Published: (2025)
Parasite: A Steganography-based Backdoor Attack Framework for Diffusion Models
by: Chen, Jiahao, et al.
Published: (2025)
by: Chen, Jiahao, et al.
Published: (2025)
CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis with Multimodal Diffusion
by: Huang, Nisha, et al.
Published: (2024)
by: Huang, Nisha, et al.
Published: (2024)
Cross-Task Attack: A Self-Supervision Generative Framework Based on Attention Shift
by: Zeng, Qingyuan, et al.
Published: (2024)
by: Zeng, Qingyuan, et al.
Published: (2024)
Diffusion-Guided Semantic Consistency for Multimodal Heterogeneity
by: Liu, Jing, et al.
Published: (2026)
by: Liu, Jing, et al.
Published: (2026)
FlexPose: Pose Distribution Adaptation with Limited Guidance
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
by: Yang, Shengzhu, et al.
Published: (2025)
by: Yang, Shengzhu, et al.
Published: (2025)
Semantic Generative Tuning for Unified Multimodal Models
by: Yu, Songsong, et al.
Published: (2026)
by: Yu, Songsong, et al.
Published: (2026)
FM-OSD: Foundation Model-Enabled One-Shot Detection of Anatomical Landmarks
by: Miao, Juzheng, et al.
Published: (2024)
by: Miao, Juzheng, et al.
Published: (2024)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
by: Lei, Zeyu, et al.
Published: (2025)
by: Lei, Zeyu, et al.
Published: (2025)
DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation
by: Wang, Zhechao, et al.
Published: (2026)
by: Wang, Zhechao, et al.
Published: (2026)
A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
DynamicID: Zero-Shot Multi-ID Image Personalization with Flexible Facial Editability
by: Hu, Xirui, et al.
Published: (2025)
by: Hu, Xirui, et al.
Published: (2025)
SUM: Saliency Unification through Mamba for Visual Attention Modeling
by: Hosseini, Alireza, et al.
Published: (2024)
by: Hosseini, Alireza, et al.
Published: (2024)
Versatile Framework with Semantic and Structural guidance for Image Reconstruction from Brain Activity
by: Lu, Yizhuo, et al.
Published: (2026)
by: Lu, Yizhuo, et al.
Published: (2026)
Multimodal SAM-adapter for Semantic Segmentation
by: Curti, Iacopo, et al.
Published: (2025)
by: Curti, Iacopo, et al.
Published: (2025)
Instruction-tuned Self-Questioning Framework for Multimodal Reasoning
by: Jang, You-Won, et al.
Published: (2025)
by: Jang, You-Won, et al.
Published: (2025)
Alternative Telescopic Displacement: An Efficient Multimodal Alignment Method
by: Qin, Jiahao, et al.
Published: (2023)
by: Qin, Jiahao, et al.
Published: (2023)
Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy
by: Huang, Jiahao, et al.
Published: (2026)
by: Huang, Jiahao, et al.
Published: (2026)
Learning Generalized and Flexible Trajectory Models from Omni-Semantic Supervision
by: Zhu, Yuanshao, et al.
Published: (2025)
by: Zhu, Yuanshao, et al.
Published: (2025)
DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing
by: Lyu, Yueming, et al.
Published: (2023)
by: Lyu, Yueming, et al.
Published: (2023)
Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection
by: Zhou, Yijun, et al.
Published: (2025)
by: Zhou, Yijun, et al.
Published: (2025)
Enlighten-Your-Voice: When Multimodal Meets Zero-shot Low-light Image Enhancement
by: Zhang, Xiaofeng, et al.
Published: (2023)
by: Zhang, Xiaofeng, et al.
Published: (2023)
Swin-TUNA : A Novel PEFT Approach for Accurate Food Image Segmentation
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
MARIS: Marine Open-Vocabulary Instance Segmentation with Geometric Enhancement and Semantic Alignment
by: Li, Bingyu, et al.
Published: (2025)
by: Li, Bingyu, et al.
Published: (2025)
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
by: He, Zheqi, et al.
Published: (2025)
by: He, Zheqi, et al.
Published: (2025)
Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations Interpretability
by: Zhu, Zhiyu, et al.
Published: (2025)
by: Zhu, Zhiyu, et al.
Published: (2025)
A-MESS: Anchor based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition
by: Shen, Yaomin, et al.
Published: (2025)
by: Shen, Yaomin, et al.
Published: (2025)
Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models
by: Sun, Li, et al.
Published: (2024)
by: Sun, Li, et al.
Published: (2024)
Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
by: Chen, Zizhao, et al.
Published: (2026)
by: Chen, Zizhao, et al.
Published: (2026)
Cognition-Inspired Dual-Stream Semantic Enhancement for Vision-Based Dynamic Emotion Modeling
by: Wang, Huanzhen, et al.
Published: (2026)
by: Wang, Huanzhen, et al.
Published: (2026)
TGC-Net: A Structure-Aware and Semantically-Aligned Framework for Text-Guided Medical Image Segmentation
by: Lin, Gaoren, et al.
Published: (2025)
by: Lin, Gaoren, et al.
Published: (2025)
SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction
by: Duan, Zaipeng, et al.
Published: (2025)
by: Duan, Zaipeng, et al.
Published: (2025)
Similar Items
-
Think Bright, Diffuse Nice: Enhancing T2I-ICL via Inductive-Bias Hint Instruction and Query Contrastive Decoding
by: Ma, Zhiyong, et al.
Published: (2026) -
FlexGen: Flexible Multi-View Generation from Text and Image Inputs
by: Xu, Xinli, et al.
Published: (2024) -
FSDENet: A Frequency and Spatial Domains based Detail Enhancement Network for Remote Sensing Semantic Segmentation
by: Fu, Jiahao, et al.
Published: (2025) -
PiPa++: Towards Unification of Domain Adaptive Semantic Segmentation via Self-supervised Learning
by: Chen, Mu, et al.
Published: (2024) -
Intelligent Parsing: An Automated Parsing Framework for Extracting Design Semantics from E-commerce Creatives
by: Li, Guandong, et al.
Published: (2023)