AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oduwole, Mardiyyah, Mireku, Prince, Adebanjo, Fatimo, Olajide, Oluwatosin, Aliyu, Mahi Aminu, Novikova, Jekaterina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911221152743424
author Oduwole, Mardiyyah
Mireku, Prince
Adebanjo, Fatimo
Olajide, Oluwatosin
Aliyu, Mahi Aminu
Novikova, Jekaterina
author_facet Oduwole, Mardiyyah
Mireku, Prince
Adebanjo, Fatimo
Olajide, Oluwatosin
Aliyu, Mahi Aminu
Novikova, Jekaterina
contents Multimodal AI research has overwhelmingly focused on high-resource languages, hindering the democratization of advancements in the field. To address this, we present AfriCaption, a comprehensive framework for multilingual image captioning in 20 African languages and our contributions are threefold: (i) a curated dataset built on Flickr8k, featuring semantically aligned captions generated via a context-aware selection and translation process; (ii) a dynamic, context-preserving pipeline that ensures ongoing quality through model ensembling and adaptive substitution; and (iii) the AfriCaption model, a 0.5B parameter vision-to-text architecture that integrates SigLIP and NLLB200 for caption generation across under-represented languages. This unified framework ensures ongoing data quality and establishes the first scalable image-captioning resource for under-represented African languages, laying the groundwork for truly inclusive multimodal AI.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
Oduwole, Mardiyyah
Mireku, Prince
Adebanjo, Fatimo
Olajide, Oluwatosin
Aliyu, Mahi Aminu
Novikova, Jekaterina
Computation and Language
Artificial Intelligence
Multimodal AI research has overwhelmingly focused on high-resource languages, hindering the democratization of advancements in the field. To address this, we present AfriCaption, a comprehensive framework for multilingual image captioning in 20 African languages and our contributions are threefold: (i) a curated dataset built on Flickr8k, featuring semantically aligned captions generated via a context-aware selection and translation process; (ii) a dynamic, context-preserving pipeline that ensures ongoing quality through model ensembling and adaptive substitution; and (iii) the AfriCaption model, a 0.5B parameter vision-to-text architecture that integrates SigLIP and NLLB200 for caption generation across under-represented languages. This unified framework ensures ongoing data quality and establishes the first scalable image-captioning resource for under-represented African languages, laying the groundwork for truly inclusive multimodal AI.
title AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.17405