MM-Diff: High-Fidelity Image Personalization via Multi-Modal Condition Integration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Zhichao, Su, Qingkun, Qin, Long, Wang, Weizhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MultiDiffSense: Diffusion-Based Multi-Modal Visuo-Tactile Image Generation Conditioned on Object Shape and Contact Pose
von: Bhouri, Sirine, et al.
Veröffentlicht: (2026)
von: Bhouri, Sirine, et al.
Veröffentlicht: (2026)
DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning
von: Liu, Weimin, et al.
Veröffentlicht: (2026)
von: Liu, Weimin, et al.
Veröffentlicht: (2026)
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
von: Zhao, Haoyu, et al.
Veröffentlicht: (2023)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2023)
HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment
von: Liao, Zhichao, et al.
Veröffentlicht: (2025)
von: Liao, Zhichao, et al.
Veröffentlicht: (2025)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models
von: Velioglu, Riza, et al.
Veröffentlicht: (2024)
von: Velioglu, Riza, et al.
Veröffentlicht: (2024)
AnchorDiff: Training-Free Concept Grounding for MM-DiTs via Anchor-Based Graph Propagation
von: Zhang, Jian, et al.
Veröffentlicht: (2026)
von: Zhang, Jian, et al.
Veröffentlicht: (2026)
Toward High-Fidelity Visual Reconstruction: From EEG-Based Conditioned Generation to Joint-Modal Guided Rebuilding
von: Gong, Zhijian, et al.
Veröffentlicht: (2026)
von: Gong, Zhijian, et al.
Veröffentlicht: (2026)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
von: Wei, Lai, et al.
Veröffentlicht: (2023)
von: Wei, Lai, et al.
Veröffentlicht: (2023)
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval
von: Xu, Mingjun, et al.
Veröffentlicht: (2025)
von: Xu, Mingjun, et al.
Veröffentlicht: (2025)
MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
von: Sheludzko, Siarhei, et al.
Veröffentlicht: (2026)
von: Sheludzko, Siarhei, et al.
Veröffentlicht: (2026)
MMPB: It's Time for Multi-Modal Personalization
von: Kim, Jaeik, et al.
Veröffentlicht: (2025)
von: Kim, Jaeik, et al.
Veröffentlicht: (2025)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
von: Yan, Bei, et al.
Veröffentlicht: (2024)
von: Yan, Bei, et al.
Veröffentlicht: (2024)
Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization
von: Zhou, Hefeng, et al.
Veröffentlicht: (2026)
von: Zhou, Hefeng, et al.
Veröffentlicht: (2026)
Dual Causal Inference: Integrating Backdoor Adjustment and Instrumental Variable Learning for Medical VQA
von: Xu, Zibo, et al.
Veröffentlicht: (2026)
von: Xu, Zibo, et al.
Veröffentlicht: (2026)
Unlabeled Cross-Center Automatic Analysis for TAAD: An Integrated Framework from Segmentation to Clinical Features
von: Liu, Mengdi, et al.
Veröffentlicht: (2026)
von: Liu, Mengdi, et al.
Veröffentlicht: (2026)
DAE-Fuse: An Adaptive Discriminative Autoencoder for Multi-Modality Image Fusion
von: Guo, Yuchen, et al.
Veröffentlicht: (2024)
von: Guo, Yuchen, et al.
Veröffentlicht: (2024)
Plug-and-Play Multi-Concept Adaptive Blending for High-Fidelity Text-to-Image Synthesis
von: Woo, Young-Beom
Veröffentlicht: (2025)
von: Woo, Young-Beom
Veröffentlicht: (2025)
DiffLoRA: Generating Personalized Low-Rank Adaptation Weights with Diffusion
von: Wu, Yujia, et al.
Veröffentlicht: (2024)
von: Wu, Yujia, et al.
Veröffentlicht: (2024)
CoDiff: Conditional Diffusion Model for Collaborative 3D Object Detection
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
FedDiff: Diffusion Model Driven Federated Learning for Multi-Modal and Multi-Clients
von: Li, DaiXun, et al.
Veröffentlicht: (2023)
von: Li, DaiXun, et al.
Veröffentlicht: (2023)
MM-Point: Multi-View Information-Enhanced Multi-Modal Self-Supervised 3D Point Cloud Understanding
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024)
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
von: Ding, Yanbo, et al.
Veröffentlicht: (2024)
von: Ding, Yanbo, et al.
Veröffentlicht: (2024)
DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
von: Wang, Zitong, et al.
Veröffentlicht: (2025)
von: Wang, Zitong, et al.
Veröffentlicht: (2025)
Obtaining Optimal Spiking Neural Network in Sequence Learning via CRNN-SNN Conversion
von: Su, Jiahao, et al.
Veröffentlicht: (2024)
von: Su, Jiahao, et al.
Veröffentlicht: (2024)
Hydra-Bench: A Benchmark for Multi-Modal Leaf Wetness Sensing
von: Liu, Yimeng, et al.
Veröffentlicht: (2025)
von: Liu, Yimeng, et al.
Veröffentlicht: (2025)
Few-Shot Medical Image Segmentation with High-Fidelity Prototypes
von: Tang, Song, et al.
Veröffentlicht: (2024)
von: Tang, Song, et al.
Veröffentlicht: (2024)
Hydra: Accurate Multi-Modal Leaf Wetness Sensing with mm-Wave and Camera Fusion
von: Liu, Yimeng, et al.
Veröffentlicht: (2025)
von: Liu, Yimeng, et al.
Veröffentlicht: (2025)
MacDiff: Unified Skeleton Modeling with Masked Conditional Diffusion
von: Wu, Lehong, et al.
Veröffentlicht: (2024)
von: Wu, Lehong, et al.
Veröffentlicht: (2024)
LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
von: Girella, Federico, et al.
Veröffentlicht: (2025)
von: Girella, Federico, et al.
Veröffentlicht: (2025)
Multi-Prompt with Depth Partitioned Cross-Modal Learning
von: Tian, Yingjie, et al.
Veröffentlicht: (2023)
von: Tian, Yingjie, et al.
Veröffentlicht: (2023)
MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
High-fidelity Person-centric Subject-to-Image Synthesis
von: Wang, Yibin, et al.
Veröffentlicht: (2023)
von: Wang, Yibin, et al.
Veröffentlicht: (2023)
Adaptive Stereo Depth Estimation with Multi-Spectral Images Across All Lighting Conditions
von: Qin, Zihan, et al.
Veröffentlicht: (2024)
von: Qin, Zihan, et al.
Veröffentlicht: (2024)
CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute Editing
von: Xian, Xiaole, et al.
Veröffentlicht: (2024)
von: Xian, Xiaole, et al.
Veröffentlicht: (2024)
Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
von: Chen, Yiwen, et al.
Veröffentlicht: (2025)
von: Chen, Yiwen, et al.
Veröffentlicht: (2025)
ColoDiff: Integrating Dynamic Consistency With Content Awareness for Colonoscopy Video Generation
von: Fu, Junhu, et al.
Veröffentlicht: (2026)
von: Fu, Junhu, et al.
Veröffentlicht: (2026)
Image-Conditional Diffusion Transformer for Underwater Image Enhancement
von: Nie, Xingyang, et al.
Veröffentlicht: (2024)
von: Nie, Xingyang, et al.
Veröffentlicht: (2024)
Latent Bias Alignment for High-Fidelity Diffusion Inversion in Real-World Image Reconstruction and Manipulation
von: Chen, Weiming, et al.
Veröffentlicht: (2026)
von: Chen, Weiming, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MultiDiffSense: Diffusion-Based Multi-Modal Visuo-Tactile Image Generation Conditioned on Object Shape and Contact Pose
von: Bhouri, Sirine, et al.
Veröffentlicht: (2026) -
DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning
von: Liu, Weimin, et al.
Veröffentlicht: (2026) -
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
von: Zhao, Haoyu, et al.
Veröffentlicht: (2023) -
HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment
von: Liao, Zhichao, et al.
Veröffentlicht: (2025) -
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)