Saved in:
| Main Authors: | Lin, Chengde, Lu, Xijun, Chen, Guangxi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2405.08114 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DAFT-GAN: Dual Affine Transformation Generative Adversarial Network for Text-Guided Image Inpainting
by: Lee, Jihoon, et al.
Published: (2024)
by: Lee, Jihoon, et al.
Published: (2024)
StegOT: Trade-offs in Steganography via Optimal Transport
by: Lin, Chengde, et al.
Published: (2025)
by: Lin, Chengde, et al.
Published: (2025)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
by: Che, Chang, et al.
Published: (2024)
by: Che, Chang, et al.
Published: (2024)
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
by: Yuan, Yu, et al.
Published: (2024)
by: Yuan, Yu, et al.
Published: (2024)
Text-to-Image Generation Via Energy-Based CLIP
by: Ganz, Roy, et al.
Published: (2024)
by: Ganz, Roy, et al.
Published: (2024)
Image-to-Image Translation with Diffusion Transformers and CLIP-Based Image Conditioning
by: Zhu, Qiang, et al.
Published: (2025)
by: Zhu, Qiang, et al.
Published: (2025)
Progressive Image Restoration via Text-Conditioned Video Generation
by: Kang, Peng, et al.
Published: (2025)
by: Kang, Peng, et al.
Published: (2025)
SwinTextUNet: Integrating CLIP-Based Text Guidance into Swin Transformer U-Nets for Medical Image Segmentation
by: Yeafi, Ashfak, et al.
Published: (2026)
by: Yeafi, Ashfak, et al.
Published: (2026)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
by: Csizmadia, Daniel, et al.
Published: (2025)
by: Csizmadia, Daniel, et al.
Published: (2025)
SAU: A Dual-Branch Network to Enhance Long-Tailed Recognition via Generative Models
by: Li, Guangxi, et al.
Published: (2024)
by: Li, Guangxi, et al.
Published: (2024)
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
by: Zhang, Qian, et al.
Published: (2024)
by: Zhang, Qian, et al.
Published: (2024)
CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
E4C: Enhance Editability for Text-Based Image Editing by Harnessing Efficient CLIP Guidance
by: Huang, Tianrui, et al.
Published: (2024)
by: Huang, Tianrui, et al.
Published: (2024)
Exploring the Adversarial Robustness of CLIP for AI-generated Image Detection
by: De Rosa, Vincenzo, et al.
Published: (2024)
by: De Rosa, Vincenzo, et al.
Published: (2024)
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
by: Kim, Seoyeon, et al.
Published: (2023)
by: Kim, Seoyeon, et al.
Published: (2023)
Adversarial Backdoor Defense in CLIP
by: Kuang, Junhao, et al.
Published: (2024)
by: Kuang, Junhao, et al.
Published: (2024)
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)
by: Gandelsman, Yossi, et al.
Published: (2023)
ATAC: Augmentation-Based Test-Time Adversarial Correction for CLIP
by: Su, Linxiang, et al.
Published: (2025)
by: Su, Linxiang, et al.
Published: (2025)
Lung Nodule Image Synthesis Driven by Two-Stage Generative Adversarial Networks
by: Cao, Lu, et al.
Published: (2026)
by: Cao, Lu, et al.
Published: (2026)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
by: Wei, Zhixiang, et al.
Published: (2025)
by: Wei, Zhixiang, et al.
Published: (2025)
PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
by: Chen, Junsong, et al.
Published: (2023)
by: Chen, Junsong, et al.
Published: (2023)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
by: Zhang, Beichen, et al.
Published: (2024)
by: Zhang, Beichen, et al.
Published: (2024)
CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval
by: Kang, Bin, et al.
Published: (2025)
by: Kang, Bin, et al.
Published: (2025)
Contrast-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment
by: Lv, Song-Lin, et al.
Published: (2025)
by: Lv, Song-Lin, et al.
Published: (2025)
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
by: Bai, Jinbin, et al.
Published: (2024)
by: Bai, Jinbin, et al.
Published: (2024)
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
by: Basu, Samyadeep, et al.
Published: (2023)
by: Basu, Samyadeep, et al.
Published: (2023)
ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
by: Yao, Xin, et al.
Published: (2025)
by: Yao, Xin, et al.
Published: (2025)
CLIP-AGIQA: Boosting the Performance of AI-Generated Image Quality Assessment with CLIP
by: Tang, Zhenchen, et al.
Published: (2024)
by: Tang, Zhenchen, et al.
Published: (2024)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
by: Zhu, Wencheng, et al.
Published: (2025)
by: Zhu, Wencheng, et al.
Published: (2025)
CLIP-VQDiffusion : Langauge Free Training of Text To Image generation using CLIP and vector quantized diffusion model
by: Han, Seungdae, et al.
Published: (2024)
by: Han, Seungdae, et al.
Published: (2024)
DeeCLIP: A Robust and Generalizable Transformer-Based Framework for Detecting AI-Generated Images
by: Keita, Mamadou, et al.
Published: (2025)
by: Keita, Mamadou, et al.
Published: (2025)
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
by: Voronov, Anton, et al.
Published: (2024)
by: Voronov, Anton, et al.
Published: (2024)
Semantic-aware Adversarial Fine-tuning for CLIP
by: Zhang, Jiacheng, et al.
Published: (2026)
by: Zhang, Jiacheng, et al.
Published: (2026)
Benchmarking PathCLIP for Pathology Image Analysis
by: Zheng, Sunyi, et al.
Published: (2024)
by: Zheng, Sunyi, et al.
Published: (2024)
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
by: Xie, Enze, et al.
Published: (2024)
by: Xie, Enze, et al.
Published: (2024)
Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis
by: Lu, Yanzuo, et al.
Published: (2025)
by: Lu, Yanzuo, et al.
Published: (2025)
AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts
by: Liu, Yufan, et al.
Published: (2025)
by: Liu, Yufan, et al.
Published: (2025)
NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
by: Yuan, Yu, et al.
Published: (2025)
by: Yuan, Yu, et al.
Published: (2025)
DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image Generation
by: Huang, Mengqi, et al.
Published: (2022)
by: Huang, Mengqi, et al.
Published: (2022)
Detecting Deepfakes with Multivariate Soft Blending and CLIP-based Image-Text Alignment
by: Li, Jingwei, et al.
Published: (2026)
by: Li, Jingwei, et al.
Published: (2026)
Similar Items
-
DAFT-GAN: Dual Affine Transformation Generative Adversarial Network for Text-Guided Image Inpainting
by: Lee, Jihoon, et al.
Published: (2024) -
StegOT: Trade-offs in Steganography via Optimal Transport
by: Lin, Chengde, et al.
Published: (2025) -
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
by: Che, Chang, et al.
Published: (2024) -
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
by: Yuan, Yu, et al.
Published: (2024) -
Text-to-Image Generation Via Energy-Based CLIP
by: Ganz, Roy, et al.
Published: (2024)