TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias
Fuente:
arXiv
Saved in:
| Main Authors: | Jo, Sanghyun, Ryu, Soohyun, Kim, Sungyub, Yang, Eunho, Kim, Kyungsu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing
by: Kim, Joowon, et al.
Published: (2025)
by: Kim, Joowon, et al.
Published: (2025)
LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models
by: Park, Sihwan, et al.
Published: (2025)
by: Park, Sihwan, et al.
Published: (2025)
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing
by: Kim, Yoonjeon, et al.
Published: (2024)
by: Kim, Yoonjeon, et al.
Published: (2024)
DHR: Dual Features-Driven Hierarchical Rebalancing in Inter- and Intra-Class Regions for Weakly-Supervised Semantic Segmentation
by: Jo, Sanghyun, et al.
Published: (2024)
by: Jo, Sanghyun, et al.
Published: (2024)
ISAC: Training-Free Instance-to-Semantic Attention Control for Improving Multi-Instance Generation
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
Tag2Text: Guiding Vision-Language Model via Image Tagging
by: Huang, Xinyu, et al.
Published: (2023)
by: Huang, Xinyu, et al.
Published: (2023)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
by: Kim, Sohee, et al.
Published: (2025)
by: Kim, Sohee, et al.
Published: (2025)
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
by: Kim, Seoyeon, et al.
Published: (2023)
by: Kim, Seoyeon, et al.
Published: (2023)
COIN: Confidence Score-Guided Distillation for Annotation-Free Cell Segmentation
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
CLIP the Landscape: Automated Tagging of Crowdsourced Landscape Images
by: Ilyankou, Ilya, et al.
Published: (2025)
by: Ilyankou, Ilya, et al.
Published: (2025)
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
OTTER: Open-Tagging via Text-Image Representation for Multi-modal Understanding
by: Ouyang, Jieer, et al.
Published: (2025)
by: Ouyang, Jieer, et al.
Published: (2025)
TRACE: Your Diffusion Model is Secretly an Instance Edge Detector
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation
by: Jo, Sanghyun, et al.
Published: (2026)
by: Jo, Sanghyun, et al.
Published: (2026)
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
OFF-CLIP: Improving Normal Detection Confidence in Radiology CLIP with Simple Off-Diagonal Term Auto-Adjustment
by: Park, Junhyun, et al.
Published: (2025)
by: Park, Junhyun, et al.
Published: (2025)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
by: Csizmadia, Daniel, et al.
Published: (2025)
by: Csizmadia, Daniel, et al.
Published: (2025)
TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification
by: Liu, Qinying, et al.
Published: (2023)
by: Liu, Qinying, et al.
Published: (2023)
CLIP-VQDiffusion : Langauge Free Training of Text To Image generation using CLIP and vector quantized diffusion model
by: Han, Seungdae, et al.
Published: (2024)
by: Han, Seungdae, et al.
Published: (2024)
TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic Segmentation
by: Li, Jingyao, et al.
Published: (2023)
by: Li, Jingyao, et al.
Published: (2023)
Data-Efficient Unsupervised Interpolation Without Any Intermediate Frame for 4D Medical Images
by: Kim, JungEun, et al.
Published: (2024)
by: Kim, JungEun, et al.
Published: (2024)
Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps
by: Kim, Jeeyung, et al.
Published: (2024)
by: Kim, Jeeyung, et al.
Published: (2024)
Safeguard Text-to-Image Diffusion Models with Human Feedback Inversion
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
Detecting Deepfakes with Multivariate Soft Blending and CLIP-based Image-Text Alignment
by: Li, Jingwei, et al.
Published: (2026)
by: Li, Jingwei, et al.
Published: (2026)
ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction Data
by: Shen, Yufan, et al.
Published: (2024)
by: Shen, Yufan, et al.
Published: (2024)
Contrast-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment
by: Lv, Song-Lin, et al.
Published: (2025)
by: Lv, Song-Lin, et al.
Published: (2025)
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
by: Basu, Samyadeep, et al.
Published: (2023)
by: Basu, Samyadeep, et al.
Published: (2023)
Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
by: Kwon, Jihoon, et al.
Published: (2025)
by: Kwon, Jihoon, et al.
Published: (2025)
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
by: Oh, Yoonjin, et al.
Published: (2025)
by: Oh, Yoonjin, et al.
Published: (2025)
Towards Scalable Human-aligned Benchmark for Text-guided Image Editing
by: Ryu, Suho, et al.
Published: (2025)
by: Ryu, Suho, et al.
Published: (2025)
Semantic Anchoring for Robust Personalization in Text-to-Image Diffusion Models
by: Yang, Seoyun, et al.
Published: (2025)
by: Yang, Seoyun, et al.
Published: (2025)
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
by: Das, Anurag, et al.
Published: (2024)
by: Das, Anurag, et al.
Published: (2024)
A Simple Remedy for Dataset Bias via Self-Influence: A Mislabeled Sample Perspective
by: Jung, Yeonsung, et al.
Published: (2024)
by: Jung, Yeonsung, et al.
Published: (2024)
Restoring Initial Noise Sensitivity in Text-to-Image Distillation via Geometric Alignment
by: Huang, Huayang, et al.
Published: (2026)
by: Huang, Huayang, et al.
Published: (2026)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
by: Park, NaHyeon, et al.
Published: (2024)
by: Park, NaHyeon, et al.
Published: (2024)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
FolkTalent: Enhancing Classification and Tagging of Indian Folk Paintings
by: Hada, Nancy, et al.
Published: (2024)
by: Hada, Nancy, et al.
Published: (2024)
Are Multimodal Large Language Models Good Annotators for Image Tagging?
by: Xie, Ming-Kun, et al.
Published: (2026)
by: Xie, Ming-Kun, et al.
Published: (2026)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
by: Che, Chang, et al.
Published: (2024)
by: Che, Chang, et al.
Published: (2024)
Prompt Augmentation for Self-supervised Text-guided Image Manipulation
by: Bodur, Rumeysa, et al.
Published: (2024)
by: Bodur, Rumeysa, et al.
Published: (2024)
Similar Items
-
Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing
by: Kim, Joowon, et al.
Published: (2025) -
LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models
by: Park, Sihwan, et al.
Published: (2025) -
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing
by: Kim, Yoonjeon, et al.
Published: (2024) -
DHR: Dual Features-Driven Hierarchical Rebalancing in Inter- and Intra-Class Regions for Weakly-Supervised Semantic Segmentation
by: Jo, Sanghyun, et al.
Published: (2024) -
ISAC: Training-Free Instance-to-Semantic Attention Control for Improving Multi-Instance Generation
by: Jo, Sanghyun, et al.
Published: (2025)