From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Chaochao, Qian, Chen, Zheng, Guodong, Fan, Hongxing, Gao, Hongzhi, Zhang, Jie, Shao, Jing, Deng, Jingyi, Fu, Jinlan, Huang, Kexin, Li, Kunchang, Li, Lijun, Wang, Limin, Sheng, Lu, Chen, Meiqi, Zhang, Ming, Ren, Qibing, Chen, Sirui, Gui, Tao, Ouyang, Wanli, Wang, Yali, Teng, Yan, Wang, Yaru, Wang, Yi, He, Yinan, Wang, Yingchun, Wang, Yixu, Zhang, Yongting, Qiao, Yu, Shen, Yujiong, Mou, Yurong, Chen, Yuxi, Zhang, Zaibin, Shi, Zhelun, Yin, Zhenfei, Wang, Zhipin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917577367748608
author Lu, Chaochao
Qian, Chen
Zheng, Guodong
Fan, Hongxing
Gao, Hongzhi
Zhang, Jie
Shao, Jing
Deng, Jingyi
Fu, Jinlan
Huang, Kexin
Li, Kunchang
Li, Lijun
Wang, Limin
Sheng, Lu
Chen, Meiqi
Zhang, Ming
Ren, Qibing
Chen, Sirui
Gui, Tao
Ouyang, Wanli
Wang, Yali
Teng, Yan
Wang, Yaru
Wang, Yi
He, Yinan
Wang, Yingchun
Wang, Yixu
Zhang, Yongting
Qiao, Yu
Shen, Yujiong
Mou, Yurong
Chen, Yuxi
Zhang, Zaibin
Shi, Zhelun
Yin, Zhenfei
Wang, Zhipin
author_facet Lu, Chaochao
Qian, Chen
Zheng, Guodong
Fan, Hongxing
Gao, Hongzhi
Zhang, Jie
Shao, Jing
Deng, Jingyi
Fu, Jinlan
Huang, Kexin
Li, Kunchang
Li, Lijun
Wang, Limin
Sheng, Lu
Chen, Meiqi
Zhang, Ming
Ren, Qibing
Chen, Sirui
Gui, Tao
Ouyang, Wanli
Wang, Yali
Teng, Yan
Wang, Yaru
Wang, Yi
He, Yinan
Wang, Yingchun
Wang, Yixu
Zhang, Yongting
Qiao, Yu
Shen, Yujiong
Mou, Yurong
Chen, Yuxi
Zhang, Zaibin
Shi, Zhelun
Yin, Zhenfei
Wang, Zhipin
contents Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents. However, there is still a wide gap between the performance of recent MLLM-based applications and the expectation of the broad public, even though the most powerful OpenAI's GPT-4 and Google's Gemini have been deployed. This paper strives to enhance understanding of the gap through the lens of a qualitative study on the generalizability, trustworthiness, and causal reasoning capabilities of recent proprietary and open-source MLLMs across four modalities: ie, text, code, image, and video, ultimately aiming to improve the transparency of MLLMs. We believe these properties are several representative factors that define the reliability of MLLMs, in supporting various downstream applications. To be specific, we evaluate the closed-source GPT-4 and Gemini and 6 open-source LLMs and MLLMs. Overall we evaluate 230 manually designed cases, where the qualitative results are then summarized into 12 scores (ie, 4 modalities times 3 properties). In total, we uncover 14 empirical findings that are useful to understand the capabilities and limitations of both proprietary and open-source MLLMs, towards more reliable downstream multi-modal applications.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15071
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
Lu, Chaochao
Qian, Chen
Zheng, Guodong
Fan, Hongxing
Gao, Hongzhi
Zhang, Jie
Shao, Jing
Deng, Jingyi
Fu, Jinlan
Huang, Kexin
Li, Kunchang
Li, Lijun
Wang, Limin
Sheng, Lu
Chen, Meiqi
Zhang, Ming
Ren, Qibing
Chen, Sirui
Gui, Tao
Ouyang, Wanli
Wang, Yali
Teng, Yan
Wang, Yaru
Wang, Yi
He, Yinan
Wang, Yingchun
Wang, Yixu
Zhang, Yongting
Qiao, Yu
Shen, Yujiong
Mou, Yurong
Chen, Yuxi
Zhang, Zaibin
Shi, Zhelun
Yin, Zhenfei
Wang, Zhipin
Computer Vision and Pattern Recognition
Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents. However, there is still a wide gap between the performance of recent MLLM-based applications and the expectation of the broad public, even though the most powerful OpenAI's GPT-4 and Google's Gemini have been deployed. This paper strives to enhance understanding of the gap through the lens of a qualitative study on the generalizability, trustworthiness, and causal reasoning capabilities of recent proprietary and open-source MLLMs across four modalities: ie, text, code, image, and video, ultimately aiming to improve the transparency of MLLMs. We believe these properties are several representative factors that define the reliability of MLLMs, in supporting various downstream applications. To be specific, we evaluate the closed-source GPT-4 and Gemini and 6 open-source LLMs and MLLMs. Overall we evaluate 230 manually designed cases, where the qualitative results are then summarized into 12 scores (ie, 4 modalities times 3 properties). In total, we uncover 14 empirical findings that are useful to understand the capabilities and limitations of both proprietary and open-source MLLMs, towards more reliable downstream multi-modal applications.
title From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.15071