Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Wenyan, Li, Jiaang, Ramos, Rita, Tang, Raphael, Elliott, Desmond |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Role of Data Curation in Image Captioning
di: Li, Wenyan, et al.
Pubblicazione: (2023)
di: Li, Wenyan, et al.
Pubblicazione: (2023)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
di: Li, Jiaang, et al.
Pubblicazione: (2025)
di: Li, Jiaang, et al.
Pubblicazione: (2025)
Towards Retrieval-Augmented Architectures for Image Captioning
di: Sarto, Sara, et al.
Pubblicazione: (2024)
di: Sarto, Sara, et al.
Pubblicazione: (2024)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
di: Fonseca, Rui, et al.
Pubblicazione: (2025)
di: Fonseca, Rui, et al.
Pubblicazione: (2025)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
di: Wu, Hao, et al.
Pubblicazione: (2024)
di: Wu, Hao, et al.
Pubblicazione: (2024)
Temporal Image Caption Retrieval Competition -- Description and Results
di: Pokrywka, Jakub, et al.
Pubblicazione: (2024)
di: Pokrywka, Jakub, et al.
Pubblicazione: (2024)
Exploring Visual Culture Awareness in GPT-4V: A Comprehensive Probing
di: Cao, Yong, et al.
Pubblicazione: (2024)
di: Cao, Yong, et al.
Pubblicazione: (2024)
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
di: Tang, Raphael, et al.
Pubblicazione: (2024)
di: Tang, Raphael, et al.
Pubblicazione: (2024)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
di: Gondal, Moazzam Umer, et al.
Pubblicazione: (2025)
di: Gondal, Moazzam Umer, et al.
Pubblicazione: (2025)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
di: Salazar, Israfel, et al.
Pubblicazione: (2025)
di: Salazar, Israfel, et al.
Pubblicazione: (2025)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
di: Huang, Yuchen, et al.
Pubblicazione: (2025)
di: Huang, Yuchen, et al.
Pubblicazione: (2025)
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
di: Yayavaram, Arnav, et al.
Pubblicazione: (2025)
di: Yayavaram, Arnav, et al.
Pubblicazione: (2025)
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
di: Dhawan, Aashish, et al.
Pubblicazione: (2026)
di: Dhawan, Aashish, et al.
Pubblicazione: (2026)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
di: Wu, Yin, et al.
Pubblicazione: (2025)
di: Wu, Yin, et al.
Pubblicazione: (2025)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
di: Gao, Sensen, et al.
Pubblicazione: (2025)
di: Gao, Sensen, et al.
Pubblicazione: (2025)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
di: Villegas, Danae Sánchez, et al.
Pubblicazione: (2025)
di: Villegas, Danae Sánchez, et al.
Pubblicazione: (2025)
Retrieval-Augmented Egocentric Video Captioning
di: Xu, Jilan, et al.
Pubblicazione: (2024)
di: Xu, Jilan, et al.
Pubblicazione: (2024)
Seeing What Tastes Good: Revisiting Multimodal Distributional Semantics in the Billion Parameter Era
di: Oneata, Dan, et al.
Pubblicazione: (2025)
di: Oneata, Dan, et al.
Pubblicazione: (2025)
Lost in Embeddings: Information Loss in Vision-Language Models
di: Li, Wenyan, et al.
Pubblicazione: (2025)
di: Li, Wenyan, et al.
Pubblicazione: (2025)
Open Vocabulary Panoptic Segmentation With Retrieval Augmentation
di: Sadeq, Nafis, et al.
Pubblicazione: (2026)
di: Sadeq, Nafis, et al.
Pubblicazione: (2026)
Towards Adaptable and Interactive Image Captioning with Data Augmentation and Episodic Memory
di: Anagnostopoulou, Aliki, et al.
Pubblicazione: (2023)
di: Anagnostopoulou, Aliki, et al.
Pubblicazione: (2023)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
di: Gao, Bingjie, et al.
Pubblicazione: (2025)
di: Gao, Bingjie, et al.
Pubblicazione: (2025)
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
VERA: Identifying and Leveraging Visual Evidence Retrieval Heads in Long-Context Understanding
di: Pei, Rongcan, et al.
Pubblicazione: (2026)
di: Pei, Rongcan, et al.
Pubblicazione: (2026)
EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
di: Li, Jiaxuan, et al.
Pubblicazione: (2023)
di: Li, Jiaxuan, et al.
Pubblicazione: (2023)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
di: Shih, Yu-Fei, et al.
Pubblicazione: (2025)
di: Shih, Yu-Fei, et al.
Pubblicazione: (2025)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
di: Wang, Zhengren, et al.
Pubblicazione: (2026)
di: Wang, Zhengren, et al.
Pubblicazione: (2026)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
di: Sun, Yubo, et al.
Pubblicazione: (2025)
di: Sun, Yubo, et al.
Pubblicazione: (2025)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
di: Gwilliam, Matthew, et al.
Pubblicazione: (2023)
di: Gwilliam, Matthew, et al.
Pubblicazione: (2023)
Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage
di: Cioni, Dario, et al.
Pubblicazione: (2023)
di: Cioni, Dario, et al.
Pubblicazione: (2023)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
di: Loo, Gowen, et al.
Pubblicazione: (2025)
di: Loo, Gowen, et al.
Pubblicazione: (2025)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
di: Kim, Taewhan, et al.
Pubblicazione: (2024)
di: Kim, Taewhan, et al.
Pubblicazione: (2024)
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
di: Zhou, Li, et al.
Pubblicazione: (2025)
di: Zhou, Li, et al.
Pubblicazione: (2025)
Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval
di: Shen, Li-Cheng, et al.
Pubblicazione: (2025)
di: Shen, Li-Cheng, et al.
Pubblicazione: (2025)
Towards Text-Image Interleaved Retrieval
di: Zhang, Xin, et al.
Pubblicazione: (2025)
di: Zhang, Xin, et al.
Pubblicazione: (2025)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
di: Dong, Kuicai, et al.
Pubblicazione: (2025)
di: Dong, Kuicai, et al.
Pubblicazione: (2025)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2025)
di: Vo, Dinh-Khoi, et al.
Pubblicazione: (2025)
AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Prompt
di: Chaturvedi, Saket S., et al.
Pubblicazione: (2025)
di: Chaturvedi, Saket S., et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Role of Data Curation in Image Captioning
di: Li, Wenyan, et al.
Pubblicazione: (2023) -
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
di: Li, Jiaang, et al.
Pubblicazione: (2025) -
Towards Retrieval-Augmented Architectures for Image Captioning
di: Sarto, Sara, et al.
Pubblicazione: (2024) -
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
di: Fonseca, Rui, et al.
Pubblicazione: (2025) -
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
di: Wu, Hao, et al.
Pubblicazione: (2024)