Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Saehyung, Yu, Sangwon, Park, Junsung, Yi, Jihun, Yoon, Sungroh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
by: Park, Junsung, et al.
Published: (2025)
by: Park, Junsung, et al.
Published: (2025)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Contextualized Visual Personalization in Vision-Language Models
by: Oh, Yeongtak, et al.
Published: (2026)
by: Oh, Yeongtak, et al.
Published: (2026)
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors
by: Lee, Jonghyun, et al.
Published: (2024)
by: Lee, Jonghyun, et al.
Published: (2024)
Normality Addition via Normality Detection in Industrial Image Anomaly Detection Models
by: Yi, Jihun, et al.
Published: (2024)
by: Yi, Jihun, et al.
Published: (2024)
CKNN: Cleansed k-Nearest Neighbor for Unsupervised Video Anomaly Detection
by: Yi, Jihun, et al.
Published: (2024)
by: Yi, Jihun, et al.
Published: (2024)
On mitigating stability-plasticity dilemma in CLIP-guided image morphing via geodesic distillation loss
by: Oh, Yeongtak, et al.
Published: (2024)
by: Oh, Yeongtak, et al.
Published: (2024)
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
Unleashing Multi-Hop Reasoning Potential in Large Language Models through Repetition of Misordered Context
by: Yu, Sangwon, et al.
Published: (2024)
by: Yu, Sangwon, et al.
Published: (2024)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
by: Shin, Chaehun, et al.
Published: (2024)
by: Shin, Chaehun, et al.
Published: (2024)
Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness
by: Hu, Xin, et al.
Published: (2026)
by: Hu, Xin, et al.
Published: (2026)
STAG: Structural Test-time Alignment of Gradients for Online Adaptation
by: Shin, Juhyeon, et al.
Published: (2024)
by: Shin, Juhyeon, et al.
Published: (2024)
Text-to-Image Rectified Flow as Plug-and-Play Priors
by: Yang, Xiaofeng, et al.
Published: (2024)
by: Yang, Xiaofeng, et al.
Published: (2024)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
by: Luo, Katie, et al.
Published: (2025)
by: Luo, Katie, et al.
Published: (2025)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
by: Chen, Jiaxing, et al.
Published: (2024)
by: Chen, Jiaxing, et al.
Published: (2024)
MaxFusion: Plug&Play Multi-Modal Generation in Text-to-Image Diffusion Models
by: Nair, Nithin Gopalakrishnan, et al.
Published: (2024)
by: Nair, Nithin Gopalakrishnan, et al.
Published: (2024)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
CleanStyle: Plug-and-Play Style Conditioning Purification for Text-to-Image Stylization
by: Feng, Xiaoman, et al.
Published: (2026)
by: Feng, Xiaoman, et al.
Published: (2026)
Plug-and-Play Diffusion Distillation
by: Hsiao, Yi-Ting, et al.
Published: (2024)
by: Hsiao, Yi-Ting, et al.
Published: (2024)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
ControlDreamer: Blending Geometry and Style in Text-to-3D
by: Oh, Yeongtak, et al.
Published: (2023)
by: Oh, Yeongtak, et al.
Published: (2023)
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models
by: Zhang, Yiming, et al.
Published: (2023)
by: Zhang, Yiming, et al.
Published: (2023)
Diffusion-Based Image-to-Image Translation by Noise Correction via Prompt Interpolation
by: Lee, Junsung, et al.
Published: (2024)
by: Lee, Junsung, et al.
Published: (2024)
Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
by: Tao, Keda, et al.
Published: (2025)
by: Tao, Keda, et al.
Published: (2025)
Deep Plug-and-Play HIO Approach for Phase Retrieval
by: Isil, Cagatay, et al.
Published: (2024)
by: Isil, Cagatay, et al.
Published: (2024)
NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image
by: Jeong, Yoonwoo, et al.
Published: (2023)
by: Jeong, Yoonwoo, et al.
Published: (2023)
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
by: Oh, Yeongtak, et al.
Published: (2026)
by: Oh, Yeongtak, et al.
Published: (2026)
Improving Geometry in Sparse-View 3DGS via Reprojection-based DoF Separation
by: Kim, Yongsung, et al.
Published: (2024)
by: Kim, Yongsung, et al.
Published: (2024)
Open-Attribute Recognition for Person Retrieval: Finding People Through Distinctive and Novel Attributes
by: Park, Minjeong, et al.
Published: (2025)
by: Park, Minjeong, et al.
Published: (2025)
Derain-Agent: A Plug-and-Play Agent Framework for Rainy Image Restoration
by: Yu, Zhaocheng, et al.
Published: (2026)
by: Yu, Zhaocheng, et al.
Published: (2026)
CBNet: A Plug-and-Play Network for Segmentation-Based Scene Text Detection
by: Zhao, Xi, et al.
Published: (2022)
by: Zhao, Xi, et al.
Published: (2022)
TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
by: Baek, Kanghyun, et al.
Published: (2025)
by: Baek, Kanghyun, et al.
Published: (2025)
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models
by: Chai, Weilong, et al.
Published: (2023)
by: Chai, Weilong, et al.
Published: (2023)
A Plug-and-Play Framework for Volumetric Light-Sheet Image Reconstruction
by: Gong, Yi, et al.
Published: (2025)
by: Gong, Yi, et al.
Published: (2025)
RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
by: Oh, Yeongtak, et al.
Published: (2025)
by: Oh, Yeongtak, et al.
Published: (2025)
Trans-Adapter: A Plug-and-Play Framework for Transparent Image Inpainting
by: Dai, Yuekun, et al.
Published: (2025)
by: Dai, Yuekun, et al.
Published: (2025)
Similar Items
-
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
by: Park, Junsung, et al.
Published: (2025) -
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025) -
Contextualized Visual Personalization in Vision-Language Models
by: Oh, Yeongtak, et al.
Published: (2026) -
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
by: Lee, Saehyung, et al.
Published: (2024) -
Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors
by: Lee, Jonghyun, et al.
Published: (2024)