From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917577367748608 |
|---|---|
| author | Lu, Chaochao Qian, Chen Zheng, Guodong Fan, Hongxing Gao, Hongzhi Zhang, Jie Shao, Jing Deng, Jingyi Fu, Jinlan Huang, Kexin Li, Kunchang Li, Lijun Wang, Limin Sheng, Lu Chen, Meiqi Zhang, Ming Ren, Qibing Chen, Sirui Gui, Tao Ouyang, Wanli Wang, Yali Teng, Yan Wang, Yaru Wang, Yi He, Yinan Wang, Yingchun Wang, Yixu Zhang, Yongting Qiao, Yu Shen, Yujiong Mou, Yurong Chen, Yuxi Zhang, Zaibin Shi, Zhelun Yin, Zhenfei Wang, Zhipin |
| author_facet | Lu, Chaochao Qian, Chen Zheng, Guodong Fan, Hongxing Gao, Hongzhi Zhang, Jie Shao, Jing Deng, Jingyi Fu, Jinlan Huang, Kexin Li, Kunchang Li, Lijun Wang, Limin Sheng, Lu Chen, Meiqi Zhang, Ming Ren, Qibing Chen, Sirui Gui, Tao Ouyang, Wanli Wang, Yali Teng, Yan Wang, Yaru Wang, Yi He, Yinan Wang, Yingchun Wang, Yixu Zhang, Yongting Qiao, Yu Shen, Yujiong Mou, Yurong Chen, Yuxi Zhang, Zaibin Shi, Zhelun Yin, Zhenfei Wang, Zhipin |
| contents | Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents. However, there is still a wide gap between the performance of recent MLLM-based applications and the expectation of the broad public, even though the most powerful OpenAI's GPT-4 and Google's Gemini have been deployed. This paper strives to enhance understanding of the gap through the lens of a qualitative study on the generalizability, trustworthiness, and causal reasoning capabilities of recent proprietary and open-source MLLMs across four modalities: ie, text, code, image, and video, ultimately aiming to improve the transparency of MLLMs. We believe these properties are several representative factors that define the reliability of MLLMs, in supporting various downstream applications. To be specific, we evaluate the closed-source GPT-4 and Gemini and 6 open-source LLMs and MLLMs. Overall we evaluate 230 manually designed cases, where the qualitative results are then summarized into 12 scores (ie, 4 modalities times 3 properties). In total, we uncover 14 empirical findings that are useful to understand the capabilities and limitations of both proprietary and open-source MLLMs, towards more reliable downstream multi-modal applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_15071 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities Lu, Chaochao Qian, Chen Zheng, Guodong Fan, Hongxing Gao, Hongzhi Zhang, Jie Shao, Jing Deng, Jingyi Fu, Jinlan Huang, Kexin Li, Kunchang Li, Lijun Wang, Limin Sheng, Lu Chen, Meiqi Zhang, Ming Ren, Qibing Chen, Sirui Gui, Tao Ouyang, Wanli Wang, Yali Teng, Yan Wang, Yaru Wang, Yi He, Yinan Wang, Yingchun Wang, Yixu Zhang, Yongting Qiao, Yu Shen, Yujiong Mou, Yurong Chen, Yuxi Zhang, Zaibin Shi, Zhelun Yin, Zhenfei Wang, Zhipin Computer Vision and Pattern Recognition Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents. However, there is still a wide gap between the performance of recent MLLM-based applications and the expectation of the broad public, even though the most powerful OpenAI's GPT-4 and Google's Gemini have been deployed. This paper strives to enhance understanding of the gap through the lens of a qualitative study on the generalizability, trustworthiness, and causal reasoning capabilities of recent proprietary and open-source MLLMs across four modalities: ie, text, code, image, and video, ultimately aiming to improve the transparency of MLLMs. We believe these properties are several representative factors that define the reliability of MLLMs, in supporting various downstream applications. To be specific, we evaluate the closed-source GPT-4 and Gemini and 6 open-source LLMs and MLLMs. Overall we evaluate 230 manually designed cases, where the qualitative results are then summarized into 12 scores (ie, 4 modalities times 3 properties). In total, we uncover 14 empirical findings that are useful to understand the capabilities and limitations of both proprietary and open-source MLLMs, towards more reliable downstream multi-modal applications. |
| title | From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2401.15071 |