A Novel Evaluation Framework for Image2Text Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Jia-Hong, Zhu, Hongyi, Shen, Yixian, Rudinac, Stevan, Pacces, Alessio M., Kanoulas, Evangelos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models
von: Zhu, Hongyi, et al.
Veröffentlicht: (2024)
von: Zhu, Hongyi, et al.
Veröffentlicht: (2024)
Personalized Video Summarization using Text-Based Queries and Conditional Modeling
von: Huang, Jia-Hong
Veröffentlicht: (2024)
von: Huang, Jia-Hong
Veröffentlicht: (2024)
Offline Evaluation of Set-Based Text-to-Image Generation
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
A Resource-Efficient Training Framework for Remote Sensing Text--Image Retrieval
von: Zhang, Weihang, et al.
Veröffentlicht: (2025)
von: Zhang, Weihang, et al.
Veröffentlicht: (2025)
ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
von: Zhang, Qi, et al.
Veröffentlicht: (2025)
von: Zhang, Qi, et al.
Veröffentlicht: (2025)
An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval
von: Byun, Jaeseok, et al.
Veröffentlicht: (2024)
von: Byun, Jaeseok, et al.
Veröffentlicht: (2024)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
von: Yang, Jiashu, et al.
Veröffentlicht: (2026)
von: Yang, Jiashu, et al.
Veröffentlicht: (2026)
ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval
von: Zhang, Zhuocheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhuocheng, et al.
Veröffentlicht: (2026)
Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated Images
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
ProGEO: Generating Prompts through Image-Text Contrastive Learning for Visual Geo-localization
von: Mao, Chen, et al.
Veröffentlicht: (2024)
von: Mao, Chen, et al.
Veröffentlicht: (2024)
ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization
von: Guo, Yuanhe, et al.
Veröffentlicht: (2025)
von: Guo, Yuanhe, et al.
Veröffentlicht: (2025)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
DEMO: A Statistical Perspective for Efficient Image-Text Matching
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models
von: Xu, Yexing, et al.
Veröffentlicht: (2026)
von: Xu, Yexing, et al.
Veröffentlicht: (2026)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
von: Le, Hoang-Bao, et al.
Veröffentlicht: (2025)
von: Le, Hoang-Bao, et al.
Veröffentlicht: (2025)
Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching
von: Ge, Xuri, et al.
Veröffentlicht: (2024)
von: Ge, Xuri, et al.
Veröffentlicht: (2024)
DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition
von: Xu, Yiyan, et al.
Veröffentlicht: (2025)
von: Xu, Yiyan, et al.
Veröffentlicht: (2025)
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
von: Tu, Rong-Cheng, et al.
Veröffentlicht: (2025)
von: Tu, Rong-Cheng, et al.
Veröffentlicht: (2025)
Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text Matching
von: Ma, Xiang, et al.
Veröffentlicht: (2024)
von: Ma, Xiang, et al.
Veröffentlicht: (2024)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
von: Sun, Zengbao, et al.
Veröffentlicht: (2024)
von: Sun, Zengbao, et al.
Veröffentlicht: (2024)
A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2025)
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2025)
Capability-aware Prompt Reformulation Learning for Text-to-Image Generation
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)
Towards Fine-Grained Citation Evaluation in Generated Text: A Comparative Analysis of Faithfulness Metrics
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
A Flexible and Scalable Framework for Video Moment Search
von: Zhang, Chongzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Chongzhi, et al.
Veröffentlicht: (2025)
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
von: Deng, Chenlong, et al.
Veröffentlicht: (2026)
von: Deng, Chenlong, et al.
Veröffentlicht: (2026)
RAGAR: Retrieval Augmented Personalized Image Generation Guided by Recommendation
von: Ling, Run, et al.
Veröffentlicht: (2025)
von: Ling, Run, et al.
Veröffentlicht: (2025)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
von: Ning, Hailong, et al.
Veröffentlicht: (2025)
von: Ning, Hailong, et al.
Veröffentlicht: (2025)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
von: Gao, Haowen, et al.
Veröffentlicht: (2025)
von: Gao, Haowen, et al.
Veröffentlicht: (2025)
The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
von: Rossetto, Luca, et al.
Veröffentlicht: (2025)
von: Rossetto, Luca, et al.
Veröffentlicht: (2025)
Towards Text-Image Interleaved Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset Usage
von: Zhang, Wenyi, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyi, et al.
Veröffentlicht: (2025)
Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024) -
Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models
von: Zhu, Hongyi, et al.
Veröffentlicht: (2024) -
Personalized Video Summarization using Text-Based Queries and Conditional Modeling
von: Huang, Jia-Hong
Veröffentlicht: (2024) -
Offline Evaluation of Set-Based Text-to-Image Generation
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024) -
A Resource-Efficient Training Framework for Remote Sensing Text--Image Retrieval
von: Zhang, Weihang, et al.
Veröffentlicht: (2025)