Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912788251672576 |
|---|---|
| author | Minh, Dao Sy Duy Kiet, Huynh Trung Quy, Nguyen Lam Phu Pham, Phu-Hoa Nguyen, Tran Chi |
| author_facet | Minh, Dao Sy Duy Kiet, Huynh Trung Quy, Nguyen Lam Phu Pham, Phu-Hoa Nguyen, Tran Chi |
| contents | Retrieving images from natural language descriptions is a core task at the intersection of computer vision and natural language processing, with wide-ranging applications in search engines, media archiving, and digital content management. However, real-world image-text retrieval remains challenging due to vague or context-dependent queries, linguistic variability, and the need for scalable solutions. In this work, we propose a lightweight two-stage retrieval pipeline that leverages event-centric entity extraction to incorporate temporal and contextual signals from real-world captions. The first stage performs efficient candidate filtering using BM25 based on salient entities, while the second stage applies BEiT-3 models to capture deep multimodal semantics and rerank the results. Evaluated on the OpenEvents v1 benchmark, our method achieves a mean average precision of 0.559, substantially outperforming prior baselines. These results highlight the effectiveness of combining event-guided filtering with long-text vision-language modeling for accurate and efficient retrieval in complex, real-world scenarios. Our code is available at https://github.com/PhamPhuHoa-23/Event-Based-Image-Retrieval |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_21221 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval Minh, Dao Sy Duy Kiet, Huynh Trung Quy, Nguyen Lam Phu Pham, Phu-Hoa Nguyen, Tran Chi Computer Vision and Pattern Recognition Artificial Intelligence I.2.10; H.3.3 Retrieving images from natural language descriptions is a core task at the intersection of computer vision and natural language processing, with wide-ranging applications in search engines, media archiving, and digital content management. However, real-world image-text retrieval remains challenging due to vague or context-dependent queries, linguistic variability, and the need for scalable solutions. In this work, we propose a lightweight two-stage retrieval pipeline that leverages event-centric entity extraction to incorporate temporal and contextual signals from real-world captions. The first stage performs efficient candidate filtering using BM25 based on salient entities, while the second stage applies BEiT-3 models to capture deep multimodal semantics and rerank the results. Evaluated on the OpenEvents v1 benchmark, our method achieves a mean average precision of 0.559, substantially outperforming prior baselines. These results highlight the effectiveness of combining event-guided filtering with long-text vision-language modeling for accurate and efficient retrieval in complex, real-world scenarios. Our code is available at https://github.com/PhamPhuHoa-23/Event-Based-Image-Retrieval |
| title | Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence I.2.10; H.3.3 |
| url | https://arxiv.org/abs/2512.21221 |