Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Deng, Shijian, Zhao, Wentian, Li, Yu-Jhe, Wan, Kun, Miranda, Daniel, Kale, Ajinkya, Tian, Yapeng
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917849595904000
author Deng, Shijian
Zhao, Wentian
Li, Yu-Jhe
Wan, Kun
Miranda, Daniel
Kale, Ajinkya
Tian, Yapeng
author_facet Deng, Shijian
Zhao, Wentian
Li, Yu-Jhe
Wan, Kun
Miranda, Daniel
Kale, Ajinkya
Tian, Yapeng
contents Self-improvement in multimodal large language models (MLLMs) is crucial for enhancing their reliability and robustness. However, current methods often rely heavily on MLLMs themselves as judges, leading to high computational costs and potential pitfalls like reward hacking and model collapse. This paper introduces a novel, model-level judge-free self-improvement framework. Our approach employs a controlled feedback mechanism while eliminating the need for MLLMs in the verification loop. We generate preference learning pairs using a controllable hallucination mechanism and optimize data quality by leveraging lightweight, contrastive language-image encoders to evaluate and reverse pairs when necessary. Evaluations across public benchmarks and our newly introduced IC dataset designed to challenge hallucination control demonstrate that our model outperforms conventional techniques. We achieve superior precision and recall with significantly lower computational demands. This method offers an efficient pathway to scalable self-improvement in MLLMs, balancing performance gains with reduced resource requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17760
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
Deng, Shijian
Zhao, Wentian
Li, Yu-Jhe
Wan, Kun
Miranda, Daniel
Kale, Ajinkya
Tian, Yapeng
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Self-improvement in multimodal large language models (MLLMs) is crucial for enhancing their reliability and robustness. However, current methods often rely heavily on MLLMs themselves as judges, leading to high computational costs and potential pitfalls like reward hacking and model collapse. This paper introduces a novel, model-level judge-free self-improvement framework. Our approach employs a controlled feedback mechanism while eliminating the need for MLLMs in the verification loop. We generate preference learning pairs using a controllable hallucination mechanism and optimize data quality by leveraging lightweight, contrastive language-image encoders to evaluate and reverse pairs when necessary. Evaluations across public benchmarks and our newly introduced IC dataset designed to challenge hallucination control demonstrate that our model outperforms conventional techniques. We achieve superior precision and recall with significantly lower computational demands. This method offers an efficient pathway to scalable self-improvement in MLLMs, balancing performance gains with reduced resource requirements.
title Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.17760