LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Helff, Lukas, Friedrich, Felix, Brack, Manuel, Kersting, Kristian, Schramowski, Patrick
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916781453475840
author Helff, Lukas
Friedrich, Felix
Brack, Manuel
Kersting, Kristian
Schramowski, Patrick
author_facet Helff, Lukas
Friedrich, Felix
Brack, Manuel
Kersting, Kristian
Schramowski, Patrick
contents This paper introduces LlavaGuard, a suite of VLM-based vision safeguards that address the critical need for reliable guardrails in the era of large-scale data and models. To this end, we establish a novel open framework, describing a customizable safety taxonomy, data preprocessing, augmentation, and training setup. For teaching a VLM safeguard on safety, we further create a multimodal safety dataset with high-quality human expert annotations, where each image is labeled with a safety rating, category, and rationale. We also employ advanced augmentations to support context-specific assessments. The resulting LlavaGuard models, ranging from 0.5B to 7B, serve as a versatile tool for evaluating the safety compliance of visual content against flexible policies. In comprehensive experiments, LlavaGuard outperforms both state-of-the-art safeguards and VLMs in accuracy and in flexibly handling different policies. Additionally, we demonstrate LlavaGuard's performance in two real-world applications: large-scale dataset annotation and moderation of text-to-image models. We make our entire framework, including the dataset, model weights, and training code.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
Helff, Lukas
Friedrich, Felix
Brack, Manuel
Kersting, Kristian
Schramowski, Patrick
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
This paper introduces LlavaGuard, a suite of VLM-based vision safeguards that address the critical need for reliable guardrails in the era of large-scale data and models. To this end, we establish a novel open framework, describing a customizable safety taxonomy, data preprocessing, augmentation, and training setup. For teaching a VLM safeguard on safety, we further create a multimodal safety dataset with high-quality human expert annotations, where each image is labeled with a safety rating, category, and rationale. We also employ advanced augmentations to support context-specific assessments. The resulting LlavaGuard models, ranging from 0.5B to 7B, serve as a versatile tool for evaluating the safety compliance of visual content against flexible policies. In comprehensive experiments, LlavaGuard outperforms both state-of-the-art safeguards and VLMs in accuracy and in flexibly handling different policies. Additionally, we demonstrate LlavaGuard's performance in two real-world applications: large-scale dataset annotation and moderation of text-to-image models. We make our entire framework, including the dataset, model weights, and training code.
title LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.05113