LLaVA-Critic: Learning to Evaluate Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913717346631680 |
|---|---|
| author | Xiong, Tianyi Wang, Xiyao Guo, Dong Ye, Qinghao Fan, Haoqi Gu, Quanquan Huang, Heng Li, Chunyuan |
| author_facet | Xiong, Tianyi Wang, Xiyao Guo, Dong Ye, Qinghao Fan, Haoqi Gu, Quanquan Huang, Heng Li, Chunyuan |
| contents | We introduce LLaVA-Critic, the first open-source large multimodal model (LMM) designed as a generalist evaluator to assess performance across a wide range of multimodal tasks. LLaVA-Critic is trained using a high-quality critic instruction-following dataset that incorporates diverse evaluation criteria and scenarios. Our experiments demonstrate the model's effectiveness in two key areas: (1) LMM-as-a-Judge, where LLaVA-Critic provides reliable evaluation scores, performing on par with or surpassing GPT models on multiple evaluation benchmarks; and (2) Preference Learning, where it generates reward signals for preference learning, enhancing model alignment capabilities. This work underscores the potential of open-source LMMs in self-critique and evaluation, setting the stage for future research into scalable, superhuman alignment feedback mechanisms for LMMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_02712 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LLaVA-Critic: Learning to Evaluate Multimodal Models Xiong, Tianyi Wang, Xiyao Guo, Dong Ye, Qinghao Fan, Haoqi Gu, Quanquan Huang, Heng Li, Chunyuan Computer Vision and Pattern Recognition Computation and Language We introduce LLaVA-Critic, the first open-source large multimodal model (LMM) designed as a generalist evaluator to assess performance across a wide range of multimodal tasks. LLaVA-Critic is trained using a high-quality critic instruction-following dataset that incorporates diverse evaluation criteria and scenarios. Our experiments demonstrate the model's effectiveness in two key areas: (1) LMM-as-a-Judge, where LLaVA-Critic provides reliable evaluation scores, performing on par with or surpassing GPT models on multiple evaluation benchmarks; and (2) Preference Learning, where it generates reward signals for preference learning, enhancing model alignment capabilities. This work underscores the potential of open-source LMMs in self-critique and evaluation, setting the stage for future research into scalable, superhuman alignment feedback mechanisms for LMMs. |
| title | LLaVA-Critic: Learning to Evaluate Multimodal Models |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2410.02712 |