Capability-aware Prompt Reformulation Learning for Text-to-Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhan, Jingtao, Ai, Qingyao, Liu, Yiqun, Chen, Jia, Ma, Shaoping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Intelligence via Trial and Error
von: Zhan, Jingtao, et al.
Veröffentlicht: (2025)
von: Zhan, Jingtao, et al.
Veröffentlicht: (2025)
Prompt Refinement with Image Pivot for Text-to-Image Generation
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)
Dynamic and Parametric Retrieval-Augmented Generation
von: Su, Weihang, et al.
Veröffentlicht: (2025)
von: Su, Weihang, et al.
Veröffentlicht: (2025)
Scaling Laws For Dense Retrieval
von: Fang, Yan, et al.
Veröffentlicht: (2024)
von: Fang, Yan, et al.
Veröffentlicht: (2024)
Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models
von: Xu, Yexing, et al.
Veröffentlicht: (2026)
von: Xu, Yexing, et al.
Veröffentlicht: (2026)
ITEm: Unsupervised Image-Text Embedding Learning for eCommerce
von: Liao, Baohao, et al.
Veröffentlicht: (2023)
von: Liao, Baohao, et al.
Veröffentlicht: (2023)
ProGEO: Generating Prompts through Image-Text Contrastive Learning for Visual Geo-localization
von: Mao, Chen, et al.
Veröffentlicht: (2024)
von: Mao, Chen, et al.
Veröffentlicht: (2024)
Towards Text-Image Interleaved Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
X-Reflect: Cross-Reflection Prompting for Multimodal Recommendation
von: Lyu, Hanjia, et al.
Veröffentlicht: (2024)
von: Lyu, Hanjia, et al.
Veröffentlicht: (2024)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
A Novel Evaluation Framework for Image2Text Generation
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
Multi-Field Tool Retrieval
von: Tang, Yichen, et al.
Veröffentlicht: (2026)
von: Tang, Yichen, et al.
Veröffentlicht: (2026)
RbFT: Robust Fine-tuning for Retrieval-Augmented Generation against Retrieval Defects
von: Tu, Yiteng, et al.
Veröffentlicht: (2025)
von: Tu, Yiteng, et al.
Veröffentlicht: (2025)
DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models
von: Su, Weihang, et al.
Veröffentlicht: (2024)
von: Su, Weihang, et al.
Veröffentlicht: (2024)
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
von: Song, Tingyu, et al.
Veröffentlicht: (2026)
von: Song, Tingyu, et al.
Veröffentlicht: (2026)
Query Augmentation by Decoding Semantics from Brain Signals
von: Ye, Ziyi, et al.
Veröffentlicht: (2024)
von: Ye, Ziyi, et al.
Veröffentlicht: (2024)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
von: Shen, Wenxuan, et al.
Veröffentlicht: (2025)
von: Shen, Wenxuan, et al.
Veröffentlicht: (2025)
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
Large Language Model Informed Patent Image Retrieval
von: Lo, Hao-Cheng, et al.
Veröffentlicht: (2024)
von: Lo, Hao-Cheng, et al.
Veröffentlicht: (2024)
Indexing Multimodal Language Models for Large-scale Image Retrieval
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026)
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026)
Offline Evaluation of Set-Based Text-to-Image Generation
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
von: Guo, Hao, et al.
Veröffentlicht: (2025)
von: Guo, Hao, et al.
Veröffentlicht: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
von: Zhao, Shu, et al.
Veröffentlicht: (2025)
von: Zhao, Shu, et al.
Veröffentlicht: (2025)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
von: Yang, Jiashu, et al.
Veröffentlicht: (2026)
von: Yang, Jiashu, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines
von: Long, Xinwei, et al.
Veröffentlicht: (2025)
von: Long, Xinwei, et al.
Veröffentlicht: (2025)
CLOSP: A Unified Semantic Space for SAR, MSI, and Text in Remote Sensing
von: Cambrin, Daniele Rege, et al.
Veröffentlicht: (2025)
von: Cambrin, Daniele Rege, et al.
Veröffentlicht: (2025)
Improving Applicability of Deep Learning based Token Classification models during Training
von: Mehra, Anket, et al.
Veröffentlicht: (2025)
von: Mehra, Anket, et al.
Veröffentlicht: (2025)
Supervised Fine-Tuning or Contrastive Learning? Towards Better Multimodal LLM Reranking
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
ReinPool: Reinforcement Learning Pooling Multi-Vector Embeddings for Retrieval System
von: Cha, Sungguk, et al.
Veröffentlicht: (2026)
von: Cha, Sungguk, et al.
Veröffentlicht: (2026)
Attention Grounded Enhancement for Visual Document Retrieval
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners
von: Yao, Yunzhi, et al.
Veröffentlicht: (2025)
von: Yao, Yunzhi, et al.
Veröffentlicht: (2025)
PRE: A Peer Review Based Large Language Model Evaluator
von: Chu, Zhumin, et al.
Veröffentlicht: (2024)
von: Chu, Zhumin, et al.
Veröffentlicht: (2024)
Enhancing Judgment Document Generation via Agentic Legal Information Collection and Rubric-Guided Optimization
von: Su, Weihang, et al.
Veröffentlicht: (2026)
von: Su, Weihang, et al.
Veröffentlicht: (2026)
Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text Matching
von: Ma, Xiang, et al.
Veröffentlicht: (2024)
von: Ma, Xiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Evaluating Intelligence via Trial and Error
von: Zhan, Jingtao, et al.
Veröffentlicht: (2025) -
Prompt Refinement with Image Pivot for Text-to-Image Generation
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024) -
Dynamic and Parametric Retrieval-Augmented Generation
von: Su, Weihang, et al.
Veröffentlicht: (2025) -
Scaling Laws For Dense Retrieval
von: Fang, Yan, et al.
Veröffentlicht: (2024) -
Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models
von: Xu, Yexing, et al.
Veröffentlicht: (2026)