MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Donghao, Huang, Jiancheng, Bai, Jinbin, Wang, Jiaze, Chen, Hao, Chen, Guangyong, Hu, Xiaowei, Heng, Pheng-Ann |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
by: Wang, Jiaze, et al.
Published: (2024)
by: Wang, Jiaze, et al.
Published: (2024)
Point Cloud Understanding via Attention-Driven Contrastive Learning
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
Distribution-Aware Calibration for Object Detection with Noisy Bounding Boxes
by: Zhou, Donghao, et al.
Published: (2023)
by: Zhou, Donghao, et al.
Published: (2023)
Unifying Physically-Informed Weather Priors in A Single Model for Image Restoration Across Multiple Adverse Weather Conditions
by: Xu, Jiaqi, et al.
Published: (2026)
by: Xu, Jiaqi, et al.
Published: (2026)
SFANet: Spatial-Frequency Attention Network for Weather Forecasting
by: Wang, Jiaze, et al.
Published: (2024)
by: Wang, Jiaze, et al.
Published: (2024)
LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs
by: Wang, Jiaze, et al.
Published: (2025)
by: Wang, Jiaze, et al.
Published: (2025)
MagicFight: Personalized Martial Arts Combat Video Generation
by: Huang, Jiancheng, et al.
Published: (2026)
by: Huang, Jiancheng, et al.
Published: (2026)
Personalized Safety Alignment for Text-to-Image Diffusion Models
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
DisCo-Layout: Disentangling and Coordinating Semantic and Physical Refinement in a Multi-Agent Framework for 3D Indoor Layout Synthesis
by: Gao, Jialin, et al.
Published: (2025)
by: Gao, Jialin, et al.
Published: (2025)
Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
by: Hu, Xiaowei, et al.
Published: (2024)
by: Hu, Xiaowei, et al.
Published: (2024)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image Editing
by: Huang, Jiancheng, et al.
Published: (2024)
by: Huang, Jiancheng, et al.
Published: (2024)
Deep Omni-supervised Learning for Rib Fracture Detection from Chest Radiology Images
by: Chai, Zhizhong, et al.
Published: (2023)
by: Chai, Zhizhong, et al.
Published: (2023)
Medical Large Vision Language Models with Multi-Image Visual Ability
by: Yang, Xikai, et al.
Published: (2025)
by: Yang, Xikai, et al.
Published: (2025)
Text-Anchored Score Composition: Tackling Condition Misalignment in Text-to-Image Diffusion Models
by: Wang, Luozhou, et al.
Published: (2023)
by: Wang, Luozhou, et al.
Published: (2023)
Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models
by: Xu, Jiaqi, et al.
Published: (2024)
by: Xu, Jiaqi, et al.
Published: (2024)
IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation
by: Zhou, Donghao, et al.
Published: (2025)
by: Zhou, Donghao, et al.
Published: (2025)
Adaptive Negative Evidential Deep Learning for Open-set Semi-supervised Learning
by: Yu, Yang, et al.
Published: (2023)
by: Yu, Yang, et al.
Published: (2023)
Towards Synchronous Memorizability and Generalizability with Site-Modulated Diffusion Replay for Cross-Site Continual Segmentation
by: Xu, Dunyuan, et al.
Published: (2024)
by: Xu, Dunyuan, et al.
Published: (2024)
Tailored Visions: Enhancing Text-to-Image Generation with Personalized Prompt Rewriting
by: Chen, Zijie, et al.
Published: (2023)
by: Chen, Zijie, et al.
Published: (2023)
Does Engram Do Memory Retrieval in Autoregressive Image Generation?
by: Wang, Jinghao, et al.
Published: (2026)
by: Wang, Jinghao, et al.
Published: (2026)
Video Instance Shadow Detection Under the Sun and Sky
by: Xing, Zhenghao, et al.
Published: (2022)
by: Xing, Zhenghao, et al.
Published: (2022)
Revisiting Shadow Detection: A New Benchmark Dataset for Complex World
by: Hu, Xiaowei, et al.
Published: (2019)
by: Hu, Xiaowei, et al.
Published: (2019)
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
by: Xing, Zhenghao, et al.
Published: (2025)
by: Xing, Zhenghao, et al.
Published: (2025)
Cross Prompting Consistency with Segment Anything Model for Semi-supervised Medical Image Segmentation
by: Miao, Juzheng, et al.
Published: (2024)
by: Miao, Juzheng, et al.
Published: (2024)
Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
Multi‐Task Mixture Density Graph Neural Networks for Predicting Catalyst Performance
by: Chen Liang, et al.
Published: (2024)
by: Chen Liang, et al.
Published: (2024)
Magic Clothing: Controllable Garment-Driven Image Synthesis
by: Chen, Weifeng, et al.
Published: (2024)
by: Chen, Weifeng, et al.
Published: (2024)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
by: Guo, Ziyu, et al.
Published: (2026)
by: Guo, Ziyu, et al.
Published: (2026)
UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation
by: Wang, Yinqiao, et al.
Published: (2025)
by: Wang, Yinqiao, et al.
Published: (2025)
SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View Adaptation
by: Wang, Yinqiao, et al.
Published: (2024)
by: Wang, Yinqiao, et al.
Published: (2024)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026)
by: Zhou, Donghao, et al.
Published: (2026)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
Component Adaptive Clustering for Generalized Category Discovery
by: Yan, Mingfu, et al.
Published: (2025)
by: Yan, Mingfu, et al.
Published: (2025)
SpecRef: A Fast Training-free Baseline of Specific Reference-Condition Real Image Editing
by: Chen, Songyan, et al.
Published: (2024)
by: Chen, Songyan, et al.
Published: (2024)
SeaDAG: Semi-autoregressive Diffusion for Conditional Directed Acyclic Graph Generation
by: Zhou, Xinyi, et al.
Published: (2024)
by: Zhou, Xinyi, et al.
Published: (2024)
MagicGeo: Training-Free Text-Guided Geometric Diagram Generation
by: Wang, Junxiao, et al.
Published: (2025)
by: Wang, Junxiao, et al.
Published: (2025)
PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval
by: Pan, Jiancheng, et al.
Published: (2024)
by: Pan, Jiancheng, et al.
Published: (2024)
Similar Items
-
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
by: Chen, Hao, et al.
Published: (2024) -
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
by: Wang, Jiaze, et al.
Published: (2024) -
Point Cloud Understanding via Attention-Driven Contrastive Learning
by: Wang, Yi, et al.
Published: (2024) -
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
by: Song, Quanjian, et al.
Published: (2025) -
Distribution-Aware Calibration for Object Detection with Noisy Bounding Boxes
by: Zhou, Donghao, et al.
Published: (2023)