Extending CLIP's Image-Text Alignment to Referring Image Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Seoyeon, Kang, Minguk, Kim, Dongwon, Park, Jaesik, Kwak, Suha |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bootstrapping Top-down Information for Self-modulating Slot Attention
by: Kim, Dongwon, et al.
Published: (2024)
by: Kim, Dongwon, et al.
Published: (2024)
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery
by: Park, Jicheol, et al.
Published: (2024)
by: Park, Jicheol, et al.
Published: (2024)
Improving Text-based Person Search via Part-level Cross-modal Correspondence
by: Park, Jicheol, et al.
Published: (2024)
by: Park, Jicheol, et al.
Published: (2024)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
by: Kim, Dongwon, et al.
Published: (2025)
by: Kim, Dongwon, et al.
Published: (2025)
Structured State-Space Regularization for Generation-Friendly Image Tokenization
by: Lee, Jinsung, et al.
Published: (2026)
by: Lee, Jinsung, et al.
Published: (2026)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026)
by: Kim, Manjin, et al.
Published: (2026)
Distilling Diffusion Models into Conditional GANs
by: Kang, Minguk, et al.
Published: (2024)
by: Kang, Minguk, et al.
Published: (2024)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025)
by: Jeong, Boseung, et al.
Published: (2025)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
by: Kim, Inho, et al.
Published: (2025)
by: Kim, Inho, et al.
Published: (2025)
FREST: Feature RESToration for Semantic Segmentation under Multiple Adverse Conditions
by: Lee, Sohyun, et al.
Published: (2024)
by: Lee, Sohyun, et al.
Published: (2024)
RePL: Pseudo-label Refinement for Semi-supervised LiDAR Semantic Segmentation
by: Kwon, Donghyeon, et al.
Published: (2026)
by: Kwon, Donghyeon, et al.
Published: (2026)
Active Label Correction for Semantic Segmentation with Foundation Models
by: Kim, Hoyoung, et al.
Published: (2024)
by: Kim, Hoyoung, et al.
Published: (2024)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
by: Kim, Dongwon, et al.
Published: (2026)
by: Kim, Dongwon, et al.
Published: (2026)
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias
by: Jo, Sanghyun, et al.
Published: (2024)
by: Jo, Sanghyun, et al.
Published: (2024)
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
by: Kim, Dongkeun, et al.
Published: (2025)
by: Kim, Dongkeun, et al.
Published: (2025)
Extreme Point Supervised Instance Segmentation
by: Lee, Hyeonjun, et al.
Published: (2024)
by: Lee, Hyeonjun, et al.
Published: (2024)
Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
by: Shin, Joonghyuk, et al.
Published: (2025)
by: Shin, Joonghyuk, et al.
Published: (2025)
Extend3D: Town-Scale 3D Generation
by: Yoon, Seungwoo, et al.
Published: (2026)
by: Yoon, Seungwoo, et al.
Published: (2026)
Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation
by: Lei, Sen, et al.
Published: (2024)
by: Lei, Sen, et al.
Published: (2024)
Improving Editability in Image Generation with Layer-wise Memory
by: Kim, Daneul, et al.
Published: (2025)
by: Kim, Daneul, et al.
Published: (2025)
VIRO: Robust and Efficient Neuro-Symbolic Reasoning with Verification for Referring Expression Comprehension
by: Park, Hyejin, et al.
Published: (2026)
by: Park, Hyejin, et al.
Published: (2026)
Towards Motion-aware Referring Image Segmentation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
CLIP-VQDiffusion : Langauge Free Training of Text To Image generation using CLIP and vector quantized diffusion model
by: Han, Seungdae, et al.
Published: (2024)
by: Han, Seungdae, et al.
Published: (2024)
Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer Learning
by: Kim, Sungyeon, et al.
Published: (2023)
by: Kim, Sungyeon, et al.
Published: (2023)
Towards More Practical Group Activity Detection: A New Benchmark and Model
by: Kim, Dongkeun, et al.
Published: (2023)
by: Kim, Dongkeun, et al.
Published: (2023)
Online Temporal Action Localization with Memory-Augmented Transformer
by: Song, Youngkil, et al.
Published: (2024)
by: Song, Youngkil, et al.
Published: (2024)
Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
by: Song, Yeji, et al.
Published: (2024)
by: Song, Yeji, et al.
Published: (2024)
Efficient and Versatile Robust Fine-Tuning of Zero-shot Models
by: Kim, Sungyeon, et al.
Published: (2024)
by: Kim, Sungyeon, et al.
Published: (2024)
ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
by: Gong, Dayoung, et al.
Published: (2024)
by: Gong, Dayoung, et al.
Published: (2024)
GaRA-SAM: Robustifying Segment Anything Model with Gated-Rank Adaptation
by: Lee, Sohyun, et al.
Published: (2025)
by: Lee, Sohyun, et al.
Published: (2025)
GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
by: Kim, Nayeong, et al.
Published: (2025)
by: Kim, Nayeong, et al.
Published: (2025)
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
by: Kim, Jimyeong, et al.
Published: (2025)
by: Kim, Jimyeong, et al.
Published: (2025)
Detecting Deepfakes with Multivariate Soft Blending and CLIP-based Image-Text Alignment
by: Li, Jingwei, et al.
Published: (2026)
by: Li, Jingwei, et al.
Published: (2026)
Leveraging Learned Image Prior for 3D Gaussian Compression
by: Shin, Seungjoo, et al.
Published: (2025)
by: Shin, Seungjoo, et al.
Published: (2025)
InstantDrag: Improving Interactivity in Drag-based Image Editing
by: Shin, Joonghyuk, et al.
Published: (2024)
by: Shin, Joonghyuk, et al.
Published: (2024)
TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
by: Lu, Qianqi, et al.
Published: (2025)
by: Lu, Qianqi, et al.
Published: (2025)
Improving Robustness to Multiple Spurious Correlations by Multi-Objective Optimization
by: Kim, Nayeong, et al.
Published: (2024)
by: Kim, Nayeong, et al.
Published: (2024)
CausalCLIPSeg: Unlocking CLIP's Potential in Referring Medical Image Segmentation with Causal Intervention
by: Chen, Yaxiong, et al.
Published: (2025)
by: Chen, Yaxiong, et al.
Published: (2025)
Distribution Matching Distillation without Fake Score Network
by: Kim, Youngjoong, et al.
Published: (2026)
by: Kim, Youngjoong, et al.
Published: (2026)
Metropolis-Hastings Sampling for 3D Gaussian Reconstruction
by: Kim, Hyunjin, et al.
Published: (2025)
by: Kim, Hyunjin, et al.
Published: (2025)
Similar Items
-
Bootstrapping Top-down Information for Self-modulating Slot Attention
by: Kim, Dongwon, et al.
Published: (2024) -
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery
by: Park, Jicheol, et al.
Published: (2024) -
Improving Text-based Person Search via Part-level Cross-modal Correspondence
by: Park, Jicheol, et al.
Published: (2024) -
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
by: Kim, Dongwon, et al.
Published: (2025) -
Structured State-Space Regularization for Generation-Friendly Image Tokenization
by: Lee, Jinsung, et al.
Published: (2026)