SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yongting, Chen, Lu, Zheng, Guodong, Gao, Yifeng, Zheng, Rui, Fu, Jinlan, Yin, Zhenfei, Jin, Senjie, Qiao, Yu, Huang, Xuanjing, Zhao, Feng, Gui, Tao, Shao, Jing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909617677664256
author Zhang, Yongting
Chen, Lu
Zheng, Guodong
Gao, Yifeng
Zheng, Rui
Fu, Jinlan
Yin, Zhenfei
Jin, Senjie
Qiao, Yu
Huang, Xuanjing
Zhao, Feng
Gui, Tao
Shao, Jing
author_facet Zhang, Yongting
Chen, Lu
Zheng, Guodong
Gao, Yifeng
Zheng, Rui
Fu, Jinlan
Yin, Zhenfei
Jin, Senjie
Qiao, Yu
Huang, Xuanjing
Zhao, Feng
Gui, Tao
Shao, Jing
contents The emergence of Vision Language Models (VLMs) has brought unprecedented advances in understanding multimodal information. The combination of textual and visual semantics in VLMs is highly complex and diverse, making the safety alignment of these models challenging. Furthermore, due to the limited study on the safety alignment of VLMs, there is a lack of large-scale, high-quality datasets. To address these limitations, we propose a Safety Preference Alignment dataset for Vision Language Models named SPA-VL. In terms of breadth, SPA-VL covers 6 harmfulness domains, 13 categories, and 53 subcategories, and contains 100,788 samples of the quadruple (question, image, chosen response, rejected response). In terms of depth, the responses are collected from 12 open-source (e.g., QwenVL) and closed-source (e.g., Gemini) VLMs to ensure diversity. The construction of preference data is fully automated, and the experimental results indicate that models trained with alignment techniques on the SPA-VL dataset exhibit substantial improvements in harmlessness and helpfulness while maintaining core capabilities. SPA-VL, as a large-scale, high-quality, and diverse dataset, represents a significant milestone in ensuring that VLMs achieve both harmlessness and helpfulness.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12030
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
Zhang, Yongting
Chen, Lu
Zheng, Guodong
Gao, Yifeng
Zheng, Rui
Fu, Jinlan
Yin, Zhenfei
Jin, Senjie
Qiao, Yu
Huang, Xuanjing
Zhao, Feng
Gui, Tao
Shao, Jing
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
The emergence of Vision Language Models (VLMs) has brought unprecedented advances in understanding multimodal information. The combination of textual and visual semantics in VLMs is highly complex and diverse, making the safety alignment of these models challenging. Furthermore, due to the limited study on the safety alignment of VLMs, there is a lack of large-scale, high-quality datasets. To address these limitations, we propose a Safety Preference Alignment dataset for Vision Language Models named SPA-VL. In terms of breadth, SPA-VL covers 6 harmfulness domains, 13 categories, and 53 subcategories, and contains 100,788 samples of the quadruple (question, image, chosen response, rejected response). In terms of depth, the responses are collected from 12 open-source (e.g., QwenVL) and closed-source (e.g., Gemini) VLMs to ensure diversity. The construction of preference data is fully automated, and the experimental results indicate that models trained with alignment techniques on the SPA-VL dataset exhibit substantial improvements in harmlessness and helpfulness while maintaining core capabilities. SPA-VL, as a large-scale, high-quality, and diverse dataset, represents a significant milestone in ensuring that VLMs achieve both harmlessness and helpfulness.
title SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2406.12030