CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Romero, David, Lyu, Chenyang, Wibowo, Haryo Akbarianto, Lynn, Teresa, Hamed, Injy, Kishore, Aditya Nanda, Mandal, Aishik, Dragonetti, Alina, Abzaliev, Artem, Tonja, Atnafu Lambebo, Balcha, Bontu Fufa, Whitehouse, Chenxi, Salamea, Christian, Velasco, Dan John, Adelani, David Ifeoluwa, Meur, David Le, Villa-Cueva, Emilio, Koto, Fajri, Farooqui, Fauzan, Belcavello, Frederico, Batnasan, Ganzorig, Vallejo, Gisela, Caulfield, Grainne, Ivetta, Guido, Song, Haiyue, Ademtew, Henok Biadglign, Maina, Hernán, Lovenia, Holy, Azime, Israel Abebe, Cruz, Jan Christian Blaise, Gala, Jay, Geng, Jiahui, Ortiz-Barajas, Jesus-German, Baek, Jinheon, Dunstan, Jocelyn, Alemany, Laura Alonso, Nagasinghe, Kumaranage Ravindu Yasas, Benotti, Luciana, D'Haro, Luis Fernando, Viridiano, Marcelo, Estecha-Garitagoitia, Marcos, Cabrera, Maria Camila Buitrago, Rodríguez-Cantelar, Mario, Jouitteau, Mélanie, Mihaylov, Mihail, Imam, Mohamed Fazli Mohamed, Adilazuarda, Muhammad Farid, Gochoo, Munkhjargal, Otgonbold, Munkh-Erdene, Etori, Naome, Niyomugisha, Olivier, Silva, Paula Mónica, Chitale, Pranjal, Dabre, Raj, Chevi, Rendi, Zhang, Ruochen, Diandaru, Ryandito, Cahyawijaya, Samuel, Góngora, Santiago, Jeong, Soyeong, Purkayastha, Sukannya, Kuribayashi, Tatsuki, Clifford, Teresa, Jayakumar, Thanmay, Torrent, Tiago Timponi, Ehsan, Toqeer, Araujo, Vladimir, Kementchedjhieva, Yova, Burzo, Zara, Lim, Zheng Wei, Yong, Zheng Xin, Ignat, Oana, Nwatu, Joan, Mihalcea, Rada, Solorio, Thamar, Aji, Alham Fikri
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912101509890048
author Romero, David
Lyu, Chenyang
Wibowo, Haryo Akbarianto
Lynn, Teresa
Hamed, Injy
Kishore, Aditya Nanda
Mandal, Aishik
Dragonetti, Alina
Abzaliev, Artem
Tonja, Atnafu Lambebo
Balcha, Bontu Fufa
Whitehouse, Chenxi
Salamea, Christian
Velasco, Dan John
Adelani, David Ifeoluwa
Meur, David Le
Villa-Cueva, Emilio
Koto, Fajri
Farooqui, Fauzan
Belcavello, Frederico
Batnasan, Ganzorig
Vallejo, Gisela
Caulfield, Grainne
Ivetta, Guido
Song, Haiyue
Ademtew, Henok Biadglign
Maina, Hernán
Lovenia, Holy
Azime, Israel Abebe
Cruz, Jan Christian Blaise
Gala, Jay
Geng, Jiahui
Ortiz-Barajas, Jesus-German
Baek, Jinheon
Dunstan, Jocelyn
Alemany, Laura Alonso
Nagasinghe, Kumaranage Ravindu Yasas
Benotti, Luciana
D'Haro, Luis Fernando
Viridiano, Marcelo
Estecha-Garitagoitia, Marcos
Cabrera, Maria Camila Buitrago
Rodríguez-Cantelar, Mario
Jouitteau, Mélanie
Mihaylov, Mihail
Imam, Mohamed Fazli Mohamed
Adilazuarda, Muhammad Farid
Gochoo, Munkhjargal
Otgonbold, Munkh-Erdene
Etori, Naome
Niyomugisha, Olivier
Silva, Paula Mónica
Chitale, Pranjal
Dabre, Raj
Chevi, Rendi
Zhang, Ruochen
Diandaru, Ryandito
Cahyawijaya, Samuel
Góngora, Santiago
Jeong, Soyeong
Purkayastha, Sukannya
Kuribayashi, Tatsuki
Clifford, Teresa
Jayakumar, Thanmay
Torrent, Tiago Timponi
Ehsan, Toqeer
Araujo, Vladimir
Kementchedjhieva, Yova
Burzo, Zara
Lim, Zheng Wei
Yong, Zheng Xin
Ignat, Oana
Nwatu, Joan
Mihalcea, Rada
Solorio, Thamar
Aji, Alham Fikri
author_facet Romero, David
Lyu, Chenyang
Wibowo, Haryo Akbarianto
Lynn, Teresa
Hamed, Injy
Kishore, Aditya Nanda
Mandal, Aishik
Dragonetti, Alina
Abzaliev, Artem
Tonja, Atnafu Lambebo
Balcha, Bontu Fufa
Whitehouse, Chenxi
Salamea, Christian
Velasco, Dan John
Adelani, David Ifeoluwa
Meur, David Le
Villa-Cueva, Emilio
Koto, Fajri
Farooqui, Fauzan
Belcavello, Frederico
Batnasan, Ganzorig
Vallejo, Gisela
Caulfield, Grainne
Ivetta, Guido
Song, Haiyue
Ademtew, Henok Biadglign
Maina, Hernán
Lovenia, Holy
Azime, Israel Abebe
Cruz, Jan Christian Blaise
Gala, Jay
Geng, Jiahui
Ortiz-Barajas, Jesus-German
Baek, Jinheon
Dunstan, Jocelyn
Alemany, Laura Alonso
Nagasinghe, Kumaranage Ravindu Yasas
Benotti, Luciana
D'Haro, Luis Fernando
Viridiano, Marcelo
Estecha-Garitagoitia, Marcos
Cabrera, Maria Camila Buitrago
Rodríguez-Cantelar, Mario
Jouitteau, Mélanie
Mihaylov, Mihail
Imam, Mohamed Fazli Mohamed
Adilazuarda, Muhammad Farid
Gochoo, Munkhjargal
Otgonbold, Munkh-Erdene
Etori, Naome
Niyomugisha, Olivier
Silva, Paula Mónica
Chitale, Pranjal
Dabre, Raj
Chevi, Rendi
Zhang, Ruochen
Diandaru, Ryandito
Cahyawijaya, Samuel
Góngora, Santiago
Jeong, Soyeong
Purkayastha, Sukannya
Kuribayashi, Tatsuki
Clifford, Teresa
Jayakumar, Thanmay
Torrent, Tiago Timponi
Ehsan, Toqeer
Araujo, Vladimir
Kementchedjhieva, Yova
Burzo, Zara
Lim, Zheng Wei
Yong, Zheng Xin
Ignat, Oana
Nwatu, Joan
Mihalcea, Rada
Solorio, Thamar
Aji, Alham Fikri
contents Visual Question Answering (VQA) is an important task in multimodal AI, and it is often used to test the ability of vision-language models to understand and reason on knowledge present in both visual and textual data. However, most of the current VQA models use datasets that are primarily focused on English and a few major world languages, with images that are typically Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, although these datasets often extend their linguistic range via translation or some other approaches, they usually keep images the same, resulting in narrow cultural representation. To address these limitations, we construct CVQA, a new Culturally-diverse multilingual Visual Question Answering benchmark, designed to cover a rich set of languages and cultures, where we engage native speakers and cultural experts in the data collection process. As a result, CVQA includes culturally-driven images and questions from across 30 countries on four continents, covering 31 languages with 13 scripts, providing a total of 10k questions. We then benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and show that the dataset is challenging for the current state-of-the-art models. This benchmark can serve as a probing evaluation suite for assessing the cultural capability and bias of multimodal models and hopefully encourage more research efforts toward increasing cultural awareness and linguistic diversity in this field.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05967
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
Romero, David
Lyu, Chenyang
Wibowo, Haryo Akbarianto
Lynn, Teresa
Hamed, Injy
Kishore, Aditya Nanda
Mandal, Aishik
Dragonetti, Alina
Abzaliev, Artem
Tonja, Atnafu Lambebo
Balcha, Bontu Fufa
Whitehouse, Chenxi
Salamea, Christian
Velasco, Dan John
Adelani, David Ifeoluwa
Meur, David Le
Villa-Cueva, Emilio
Koto, Fajri
Farooqui, Fauzan
Belcavello, Frederico
Batnasan, Ganzorig
Vallejo, Gisela
Caulfield, Grainne
Ivetta, Guido
Song, Haiyue
Ademtew, Henok Biadglign
Maina, Hernán
Lovenia, Holy
Azime, Israel Abebe
Cruz, Jan Christian Blaise
Gala, Jay
Geng, Jiahui
Ortiz-Barajas, Jesus-German
Baek, Jinheon
Dunstan, Jocelyn
Alemany, Laura Alonso
Nagasinghe, Kumaranage Ravindu Yasas
Benotti, Luciana
D'Haro, Luis Fernando
Viridiano, Marcelo
Estecha-Garitagoitia, Marcos
Cabrera, Maria Camila Buitrago
Rodríguez-Cantelar, Mario
Jouitteau, Mélanie
Mihaylov, Mihail
Imam, Mohamed Fazli Mohamed
Adilazuarda, Muhammad Farid
Gochoo, Munkhjargal
Otgonbold, Munkh-Erdene
Etori, Naome
Niyomugisha, Olivier
Silva, Paula Mónica
Chitale, Pranjal
Dabre, Raj
Chevi, Rendi
Zhang, Ruochen
Diandaru, Ryandito
Cahyawijaya, Samuel
Góngora, Santiago
Jeong, Soyeong
Purkayastha, Sukannya
Kuribayashi, Tatsuki
Clifford, Teresa
Jayakumar, Thanmay
Torrent, Tiago Timponi
Ehsan, Toqeer
Araujo, Vladimir
Kementchedjhieva, Yova
Burzo, Zara
Lim, Zheng Wei
Yong, Zheng Xin
Ignat, Oana
Nwatu, Joan
Mihalcea, Rada
Solorio, Thamar
Aji, Alham Fikri
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Visual Question Answering (VQA) is an important task in multimodal AI, and it is often used to test the ability of vision-language models to understand and reason on knowledge present in both visual and textual data. However, most of the current VQA models use datasets that are primarily focused on English and a few major world languages, with images that are typically Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, although these datasets often extend their linguistic range via translation or some other approaches, they usually keep images the same, resulting in narrow cultural representation. To address these limitations, we construct CVQA, a new Culturally-diverse multilingual Visual Question Answering benchmark, designed to cover a rich set of languages and cultures, where we engage native speakers and cultural experts in the data collection process. As a result, CVQA includes culturally-driven images and questions from across 30 countries on four continents, covering 31 languages with 13 scripts, providing a total of 10k questions. We then benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and show that the dataset is challenging for the current state-of-the-art models. This benchmark can serve as a probing evaluation suite for assessing the cultural capability and bias of multimodal models and hopefully encourage more research efforts toward increasing cultural awareness and linguistic diversity in this field.
title CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2406.05967