Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yuhao, Zhu, Zhiyuan, Liu, Heyang, Liao, Yusheng, Liu, Hongcheng, Wang, Yanfeng, Wang, Yu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916525027360768
author Wang, Yuhao
Zhu, Zhiyuan
Liu, Heyang
Liao, Yusheng
Liu, Hongcheng
Wang, Yanfeng
Wang, Yu
author_facet Wang, Yuhao
Zhu, Zhiyuan
Liu, Heyang
Liao, Yusheng
Liu, Hongcheng
Wang, Yanfeng
Wang, Yu
contents Multimodal large language models (MLLMs) excel at multimodal perception and understanding, yet their tendency to generate hallucinated or inaccurate responses undermines their trustworthiness. Existing methods have largely overlooked the importance of refusal responses as a means of enhancing MLLMs reliability. To bridge this gap, we present the Information Boundary-aware Learning Framework (InBoL), a novel approach that empowers MLLMs to refuse to answer user queries when encountering insufficient information. To the best of our knowledge, InBoL is the first framework that systematically defines the conditions under which refusal is appropriate for MLLMs using the concept of information boundaries proposed in our paper. This framework introduces a comprehensive data generation pipeline and tailored training strategies to improve the model's ability to deliver appropriate refusal responses. To evaluate the trustworthiness of MLLMs, we further propose a user-centric alignment goal along with corresponding metrics. Experimental results demonstrate a significant improvement in refusal accuracy without noticeably compromising the model's helpfulness, establishing InBoL as a pivotal advancement in building more trustworthy MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11196
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
Wang, Yuhao
Zhu, Zhiyuan
Liu, Heyang
Liao, Yusheng
Liu, Hongcheng
Wang, Yanfeng
Wang, Yu
Computation and Language
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) excel at multimodal perception and understanding, yet their tendency to generate hallucinated or inaccurate responses undermines their trustworthiness. Existing methods have largely overlooked the importance of refusal responses as a means of enhancing MLLMs reliability. To bridge this gap, we present the Information Boundary-aware Learning Framework (InBoL), a novel approach that empowers MLLMs to refuse to answer user queries when encountering insufficient information. To the best of our knowledge, InBoL is the first framework that systematically defines the conditions under which refusal is appropriate for MLLMs using the concept of information boundaries proposed in our paper. This framework introduces a comprehensive data generation pipeline and tailored training strategies to improve the model's ability to deliver appropriate refusal responses. To evaluate the trustworthiness of MLLMs, we further propose a user-centric alignment goal along with corresponding metrics. Experimental results demonstrate a significant improvement in refusal accuracy without noticeably compromising the model's helpfulness, establishing InBoL as a pivotal advancement in building more trustworthy MLLMs.
title Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.11196