IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Jiayi, Yan, Chuanhao, Xu, Xingqian, Wang, Yulin, Wang, Kai, Huang, Gao, Shi, Humphrey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
by: Xu, Xingqian, et al.
Published: (2022)
by: Xu, Xingqian, et al.
Published: (2022)
Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment
by: Guo, Jiayi, et al.
Published: (2024)
by: Guo, Jiayi, et al.
Published: (2024)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025)
by: Peruzzo, Elia, et al.
Published: (2025)
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
by: Goel, Vidit, et al.
Published: (2023)
by: Goel, Vidit, et al.
Published: (2023)
Safe Vision-Language Models via Unsafe Weights Manipulation
by: D'Incà, Moreno, et al.
Published: (2025)
by: D'Incà, Moreno, et al.
Published: (2025)
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)
by: D'Incà, Moreno, et al.
Published: (2024)
Efficient Image Generation with Variadic Attention Heads
by: Walton, Steven, et al.
Published: (2022)
by: Walton, Steven, et al.
Published: (2022)
UVMap-ID: A Controllable and Personalized UV Map Generative Model
by: Wang, Weijie, et al.
Published: (2024)
by: Wang, Weijie, et al.
Published: (2024)
ClassDiffusion: More Aligned Personalization Tuning with Explicit Class Guidance
by: Huang, Jiannan, et al.
Published: (2024)
by: Huang, Jiannan, et al.
Published: (2024)
VASE: Object-Centric Appearance and Shape Manipulation of Real Videos
by: Peruzzo, Elia, et al.
Published: (2024)
by: Peruzzo, Elia, et al.
Published: (2024)
OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
by: Su, Zhaochen, et al.
Published: (2025)
by: Su, Zhaochen, et al.
Published: (2025)
Implicit and Explicit Language Guidance for Diffusion-based Visual Perception
by: Wang, Hefeng, et al.
Published: (2024)
by: Wang, Hefeng, et al.
Published: (2024)
Meta-Semi: A Meta-learning Approach for Semi-supervised Learning
by: Wang, Yulin, et al.
Published: (2020)
by: Wang, Yulin, et al.
Published: (2020)
EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
by: Yue, Yang, et al.
Published: (2025)
by: Yue, Yang, et al.
Published: (2025)
OpenBias: Open-set Bias Detection in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)
by: D'Incà, Moreno, et al.
Published: (2024)
Implicit Concept Removal of Diffusion Models
by: Liu, Zhili, et al.
Published: (2023)
by: Liu, Zhili, et al.
Published: (2023)
Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models
by: Guo, Jiayi, et al.
Published: (2026)
by: Guo, Jiayi, et al.
Published: (2026)
Normalized Attention Guidance: Universal Negative Guidance for Diffusion Models
by: Chen, Dar-Yen, et al.
Published: (2025)
by: Chen, Dar-Yen, et al.
Published: (2025)
Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024)
by: Jain, Jitesh, et al.
Published: (2024)
CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibration
by: Guo, Xiefan, et al.
Published: (2026)
by: Guo, Xiefan, et al.
Published: (2026)
HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion Models
by: Manukyan, Hayk, et al.
Published: (2023)
by: Manukyan, Hayk, et al.
Published: (2023)
The Silent Assistant: NoiseQuery as Implicit Guidance for Goal-Driven Image Generation
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
Diffusion for Natural Image Matting
by: Hu, Yihan, et al.
Published: (2023)
by: Hu, Yihan, et al.
Published: (2023)
COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing
by: Wang, Jiangshan, et al.
Published: (2024)
by: Wang, Jiangshan, et al.
Published: (2024)
DSPFusion: Image Fusion via Degradation and Semantic Dual-Prior Guidance
by: Tang, Linfeng, et al.
Published: (2025)
by: Tang, Linfeng, et al.
Published: (2025)
LazyMAR: Accelerating Masked Autoregressive Models via Feature Caching
by: Yan, Feihong, et al.
Published: (2025)
by: Yan, Feihong, et al.
Published: (2025)
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
by: Wang, Pengfei, et al.
Published: (2025)
by: Wang, Pengfei, et al.
Published: (2025)
UniTTA: Unified Benchmark and Versatile Framework Towards Realistic Test-Time Adaptation
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Foodfusion: A Novel Approach for Food Image Composition via Diffusion Models
by: Shi, Chaohua, et al.
Published: (2024)
by: Shi, Chaohua, et al.
Published: (2024)
Seeing Through the Rain: Resolving High-Frequency Conflicts in Deraining and Super-Resolution via Diffusion Guidance
by: Li, Wenjie, et al.
Published: (2025)
by: Li, Wenjie, et al.
Published: (2025)
SARD: Segmentation-Aware Anomaly Synthesis via Region-Constrained Diffusion with Discriminative Mask Guidance
by: Wang, Yanshu, et al.
Published: (2025)
by: Wang, Yanshu, et al.
Published: (2025)
Superpixel-informed Implicit Neural Representation for Multi-Dimensional Data
by: Li, Jiayi, et al.
Published: (2024)
by: Li, Jiayi, et al.
Published: (2024)
AdaGen: Learning Adaptive Policy for Image Synthesis
by: Ni, Zanlin, et al.
Published: (2026)
by: Ni, Zanlin, et al.
Published: (2026)
Enhancing Video Super-Resolution via Implicit Resampling-based Alignment
by: Xu, Kai, et al.
Published: (2023)
by: Xu, Kai, et al.
Published: (2023)
HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models
by: Lu, Zhiguang, et al.
Published: (2025)
by: Lu, Zhiguang, et al.
Published: (2025)
Improving Diffusion Generalization with Weak-to-Strong Segmented Guidance
by: Yuan, Liangyu, et al.
Published: (2026)
by: Yuan, Liangyu, et al.
Published: (2026)
ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
by: Huang, Zhenyang, et al.
Published: (2025)
by: Huang, Zhenyang, et al.
Published: (2025)
Trajectory-Consistent Calibration for Cache-Accelerated Diffusion Models
by: Liang, Mingyu, et al.
Published: (2026)
by: Liang, Mingyu, et al.
Published: (2026)
Diffusion Models with Implicit Guidance for Medical Anomaly Detection
by: Bercea, Cosmin I., et al.
Published: (2024)
by: Bercea, Cosmin I., et al.
Published: (2024)
Similar Items
-
Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
by: Xu, Xingqian, et al.
Published: (2022) -
Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment
by: Guo, Jiayi, et al.
Published: (2024) -
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025) -
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
by: Goel, Vidit, et al.
Published: (2023) -
Safe Vision-Language Models via Unsafe Weights Manipulation
by: D'Incà, Moreno, et al.
Published: (2025)