Saved in:
| Main Authors: | Li, Sifan, Tao, Ming, Zhao, Hao, Shao, Ling, Tang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.14341 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
by: Huang, Zhe, et al.
Published: (2025)
by: Huang, Zhe, et al.
Published: (2025)
MICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept Alignment
by: Bie, Yequan, et al.
Published: (2024)
by: Bie, Yequan, et al.
Published: (2024)
GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval
by: Zou, Hao, et al.
Published: (2025)
by: Zou, Hao, et al.
Published: (2025)
Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
Uncovering the Text Embedding in Text-to-Image Diffusion Models
by: Yu, Hu, et al.
Published: (2024)
by: Yu, Hu, et al.
Published: (2024)
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)
by: Lei, Shiye, et al.
Published: (2023)
Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model
by: Chen, Hongxu, et al.
Published: (2025)
by: Chen, Hongxu, et al.
Published: (2025)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
by: Huang, Ziwei, et al.
Published: (2024)
by: Huang, Ziwei, et al.
Published: (2024)
360PanT: Training-Free Text-Driven 360-Degree Panorama-to-Panorama Translation
by: Wang, Hai, et al.
Published: (2024)
by: Wang, Hai, et al.
Published: (2024)
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
by: Xian, Jia Jun Cheng, et al.
Published: (2025)
by: Xian, Jia Jun Cheng, et al.
Published: (2025)
Concept Complement Bottleneck Model for Interpretable Medical Image Diagnosis
by: Wang, Hongmei, et al.
Published: (2024)
by: Wang, Hongmei, et al.
Published: (2024)
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
Causal-Adapter: Taming Text-to-Image Diffusion for Faithful Counterfactual Generation
by: Tong, Lei, et al.
Published: (2025)
by: Tong, Lei, et al.
Published: (2025)
4D-GSW: Kinematic-Aware Spatio-Temporal Consistent Watermarking for 4D Gaussian Splatting
by: Zhou, Sifan, et al.
Published: (2026)
by: Zhou, Sifan, et al.
Published: (2026)
Concept Conductor: Orchestrating Multiple Personalized Concepts in Text-to-Image Synthesis
by: Yao, Zebin, et al.
Published: (2024)
by: Yao, Zebin, et al.
Published: (2024)
Text Guided Image Editing with Automatic Concept Locating and Forgetting
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
by: Woo, Young Beom, et al.
Published: (2025)
by: Woo, Young Beom, et al.
Published: (2025)
Instant Preference Alignment for Text-to-Image Diffusion Models
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
by: Zhang, Yongheng, et al.
Published: (2025)
by: Zhang, Yongheng, et al.
Published: (2025)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction
by: Hao, Shaozhe, et al.
Published: (2024)
by: Hao, Shaozhe, et al.
Published: (2024)
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
by: Wang, Bingli, et al.
Published: (2026)
by: Wang, Bingli, et al.
Published: (2026)
SAM2 for Image and Video Segmentation: A Comprehensive Survey
by: Jiaxing, Zhang, et al.
Published: (2025)
by: Jiaxing, Zhang, et al.
Published: (2025)
ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion
by: Liu, Hanpeng, et al.
Published: (2026)
by: Liu, Hanpeng, et al.
Published: (2026)
Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models
by: Seo, Hoigi, et al.
Published: (2026)
by: Seo, Hoigi, et al.
Published: (2026)
GI-NAS: Boosting Gradient Inversion Attacks Through Adaptive Neural Architecture Search
by: Yu, Wenbo, et al.
Published: (2024)
by: Yu, Wenbo, et al.
Published: (2024)
Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models
by: Kwon, Gihyun, et al.
Published: (2024)
by: Kwon, Gihyun, et al.
Published: (2024)
Personalized Safety Alignment for Text-to-Image Diffusion Models
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
Life-IQA: Boosting Blind Image Quality Assessment through GCN-enhanced Layer Interaction and MoE-based Feature Decoupling
by: Tang, Long, et al.
Published: (2025)
by: Tang, Long, et al.
Published: (2025)
TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
by: Liang, Yuanzhi, et al.
Published: (2026)
by: Liang, Yuanzhi, et al.
Published: (2026)
Cycle Diffusion Model for Counterfactual Image Generation
by: Huang, Fangrui, et al.
Published: (2025)
by: Huang, Fangrui, et al.
Published: (2025)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
by: Zhang, Huixuan, et al.
Published: (2025)
by: Zhang, Huixuan, et al.
Published: (2025)
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
by: Tang, Jianting, et al.
Published: (2025)
by: Tang, Jianting, et al.
Published: (2025)
Robust Latent Matters: Boosting Image Generation with Sampling Error Synthesis
by: Qiu, Kai, et al.
Published: (2025)
by: Qiu, Kai, et al.
Published: (2025)
A Comprehensive Survey on Concept Erasure in Text-to-Image Diffusion Models
by: Kim, Changhoon, et al.
Published: (2025)
by: Kim, Changhoon, et al.
Published: (2025)
Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute Editing
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Boosting Medical Image-based Cancer Detection via Text-guided Supervision from Reports
by: Guo, Guangyu, et al.
Published: (2024)
by: Guo, Guangyu, et al.
Published: (2024)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
by: Saxon, Michael, et al.
Published: (2024)
by: Saxon, Michael, et al.
Published: (2024)
Similar Items
-
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
by: Huang, Zhe, et al.
Published: (2025) -
MICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept Alignment
by: Bie, Yequan, et al.
Published: (2024) -
GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval
by: Zou, Hao, et al.
Published: (2025) -
Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space
by: Wang, Xiaoce, et al.
Published: (2026) -
Uncovering the Text Embedding in Text-to-Image Diffusion Models
by: Yu, Hu, et al.
Published: (2024)