Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Quy, Nguyen Lam Phu, Hoa, Pham Phu, Nguyen, Tran Chi, Minh, Dao Sy Duy, Ngoc, Nguyen Hoang Minh, Kiet, Huynh Trung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917238121955328
author Quy, Nguyen Lam Phu
Hoa, Pham Phu
Nguyen, Tran Chi
Minh, Dao Sy Duy
Ngoc, Nguyen Hoang Minh
Kiet, Huynh Trung
author_facet Quy, Nguyen Lam Phu
Hoa, Pham Phu
Nguyen, Tran Chi
Minh, Dao Sy Duy
Ngoc, Nguyen Hoang Minh
Kiet, Huynh Trung
contents Real-world image captions often lack contextual depth, omitting crucial details such as event background, temporal cues, outcomes, and named entities that are not visually discernible. This gap limits the effectiveness of image understanding in domains like journalism, education, and digital archives, where richer, more informative descriptions are essential. To address this, we propose a multimodal pipeline that augments visual input with external textual knowledge. Our system retrieves semantically similar images using BEIT-3 (Flickr30k-384 and COCO-384) and SigLIP So-384, reranks them using ORB and SIFT for geometric alignment, and extracts contextual information from related articles via semantic search. A fine-tuned Qwen3 model with QLoRA then integrates this context with base captions generated by Instruct BLIP (Vicuna-7B) to produce event-enriched, context-aware descriptions. Evaluated on the OpenEvents v1 dataset, our approach generates significantly more informative captions compared to traditional methods, showing strong potential for real-world applications requiring deeper visual-textual understanding
format Preprint
id arxiv_https___arxiv_org_abs_2512_20042
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
Quy, Nguyen Lam Phu
Hoa, Pham Phu
Nguyen, Tran Chi
Minh, Dao Sy Duy
Ngoc, Nguyen Hoang Minh
Kiet, Huynh Trung
Computer Vision and Pattern Recognition
Artificial Intelligence
I.2.10; H.3.3; I.4.8
Real-world image captions often lack contextual depth, omitting crucial details such as event background, temporal cues, outcomes, and named entities that are not visually discernible. This gap limits the effectiveness of image understanding in domains like journalism, education, and digital archives, where richer, more informative descriptions are essential. To address this, we propose a multimodal pipeline that augments visual input with external textual knowledge. Our system retrieves semantically similar images using BEIT-3 (Flickr30k-384 and COCO-384) and SigLIP So-384, reranks them using ORB and SIFT for geometric alignment, and extracts contextual information from related articles via semantic search. A fine-tuned Qwen3 model with QLoRA then integrates this context with base captions generated by Instruct BLIP (Vicuna-7B) to produce event-enriched, context-aware descriptions. Evaluated on the OpenEvents v1 dataset, our approach generates significantly more informative captions compared to traditional methods, showing strong potential for real-world applications requiring deeper visual-textual understanding
title Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
topic Computer Vision and Pattern Recognition
Artificial Intelligence
I.2.10; H.3.3; I.4.8
url https://arxiv.org/abs/2512.20042