DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Zhen, Zongwei, Zeng, Biqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Iterative Prompt Refinement for Safer Text-to-Image Generation
by: Jeon, Jinwoo, et al.
Published: (2025)
by: Jeon, Jinwoo, et al.
Published: (2025)
AutoDIR: Automatic All-in-One Image Restoration with Latent Diffusion
by: Jiang, Yitong, et al.
Published: (2023)
by: Jiang, Yitong, et al.
Published: (2023)
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration
by: Zhang, Yisheng, et al.
Published: (2026)
by: Zhang, Yisheng, et al.
Published: (2026)
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
DiaLoc: An Iterative Approach to Embodied Dialog Localization
by: Zhang, Chao, et al.
Published: (2024)
by: Zhang, Chao, et al.
Published: (2024)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
by: Jeong, Suchae, et al.
Published: (2025)
by: Jeong, Suchae, et al.
Published: (2025)
Teaching Text-to-Image Models to Communicate in Dialog
by: Sun, Xiaowen, et al.
Published: (2023)
by: Sun, Xiaowen, et al.
Published: (2023)
Prompt Refinement with Image Pivot for Text-to-Image Generation
by: Zhan, Jingtao, et al.
Published: (2024)
by: Zhan, Jingtao, et al.
Published: (2024)
SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
by: Jain, Chelsi, et al.
Published: (2025)
by: Jain, Chelsi, et al.
Published: (2025)
I2PRef: Image-Driven Point Completion with Iterative Refinement
by: Hussian, Azhar, et al.
Published: (2026)
by: Hussian, Azhar, et al.
Published: (2026)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
by: Jin, Hongbo, et al.
Published: (2026)
by: Jin, Hongbo, et al.
Published: (2026)
On RGB-TIR Stereo Calibration under Extreme Resolution Asymmetry
by: Król, Michał, et al.
Published: (2026)
by: Król, Michał, et al.
Published: (2026)
DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
by: Huang, Minbin, et al.
Published: (2024)
by: Huang, Minbin, et al.
Published: (2024)
Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement
by: Guo, Junrong, et al.
Published: (2026)
by: Guo, Junrong, et al.
Published: (2026)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
by: Han, Donghoon, et al.
Published: (2024)
by: Han, Donghoon, et al.
Published: (2024)
Enhancing Image Matting in Real-World Scenes with Mask-Guided Iterative Refinement
by: Liu, Rui
Published: (2025)
by: Liu, Rui
Published: (2025)
Safe-SD: Safe and Traceable Stable Diffusion with Text Prompt Trigger for Invisible Generative Watermarking
by: Ma, Zhiyuan, et al.
Published: (2024)
by: Ma, Zhiyuan, et al.
Published: (2024)
I2I-PR: Deep Iterative Refinement for Phase Retrieval using Image-to-Image Diffusion Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
G-Refine: A General Quality Refiner for Text-to-Image Generation
by: Li, Chunyi, et al.
Published: (2024)
by: Li, Chunyi, et al.
Published: (2024)
CFR-ICL: Cascade-Forward Refinement with Iterative Click Loss for Interactive Image Segmentation
by: Sun, Shoukun, et al.
Published: (2023)
by: Sun, Shoukun, et al.
Published: (2023)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
by: Gao, Zhe, et al.
Published: (2026)
by: Gao, Zhe, et al.
Published: (2026)
Hybrid, Unified and Iterative: A Novel Framework for Text-based Person Anomaly Retrieval
by: Nguyen, Tien-Huy, et al.
Published: (2025)
by: Nguyen, Tien-Huy, et al.
Published: (2025)
Interpretable Text-Guided Image Clustering via Iterative Search
by: Zhao, Bingchen, et al.
Published: (2025)
by: Zhao, Bingchen, et al.
Published: (2025)
Improving Visual Reasoning with Iterative Evidence Refinement
by: Shi, Zeru, et al.
Published: (2026)
by: Shi, Zeru, et al.
Published: (2026)
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
by: Zhang, Haoyu, et al.
Published: (2023)
by: Zhang, Haoyu, et al.
Published: (2023)
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
by: Zeng, Ziyun, et al.
Published: (2025)
by: Zeng, Ziyun, et al.
Published: (2025)
A Vessel Bifurcation Landmark Pair Dataset for Abdominal CT Deformable Image Registration (DIR) Validation
by: Criscuolo, Edward R, et al.
Published: (2025)
by: Criscuolo, Edward R, et al.
Published: (2025)
Anatomy-Aware Conditional Image-Text Retrieval
by: Zheng, Meng, et al.
Published: (2025)
by: Zheng, Meng, et al.
Published: (2025)
Zero-shot Composed Text-Image Retrieval
by: Liu, Yikun, et al.
Published: (2023)
by: Liu, Yikun, et al.
Published: (2023)
Exploring Iterative Refinement with Diffusion Models for Video Grounding
by: Liang, Xiao, et al.
Published: (2023)
by: Liang, Xiao, et al.
Published: (2023)
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
by: Ozaki, Shintaro, et al.
Published: (2025)
by: Ozaki, Shintaro, et al.
Published: (2025)
The Solution for the GAIIC2024 RGB-TIR object detection Challenge
by: Wu, Xiangyu, et al.
Published: (2024)
by: Wu, Xiangyu, et al.
Published: (2024)
Beyond Masks: The Case for Medical Image Parsing
by: Gupta, Siddharth, et al.
Published: (2026)
by: Gupta, Siddharth, et al.
Published: (2026)
DIR-BHRNet: A Lightweight Network for Real-time Vision-based Multi-person Pose Estimation on Smartphones
by: Lan, Gongjin, et al.
Published: (2024)
by: Lan, Gongjin, et al.
Published: (2024)
Compositional Image-Text Matching and Retrieval by Grounding Entities
by: Vongala, Madhukar Reddy, et al.
Published: (2025)
by: Vongala, Madhukar Reddy, et al.
Published: (2025)
LATTE: Improving Latex Recognition for Tables and Formulae with Iterative Refinement
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
PSDiff: Diffusion Model for Person Search with Iterative and Collaborative Refinement
by: Jia, Chengyou, et al.
Published: (2023)
by: Jia, Chengyou, et al.
Published: (2023)
Similar Items
-
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024) -
Iterative Prompt Refinement for Safer Text-to-Image Generation
by: Jeon, Jinwoo, et al.
Published: (2025) -
AutoDIR: Automatic All-in-One Image Restoration with Latent Diffusion
by: Jiang, Yitong, et al.
Published: (2023) -
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration
by: Zhang, Yisheng, et al.
Published: (2026) -
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
by: Li, Ming, et al.
Published: (2025)