SafeVid: Toward Safety Aligned Video Large Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yixu, Song, Jiaxin, Gao, Yifeng, Wang, Xin, Yao, Yang, Teng, Yan, Ma, Xingjun, Wang, Yingchun, Jiang, Yu-Gang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910950737575936
author Wang, Yixu
Song, Jiaxin
Gao, Yifeng
Wang, Xin
Yao, Yang
Teng, Yan
Ma, Xingjun
Wang, Yingchun
Jiang, Yu-Gang
author_facet Wang, Yixu
Song, Jiaxin
Gao, Yifeng
Wang, Xin
Yao, Yang
Teng, Yan
Ma, Xingjun
Wang, Yingchun
Jiang, Yu-Gang
contents As Video Large Multimodal Models (VLMMs) rapidly advance, their inherent complexity introduces significant safety challenges, particularly the issue of mismatched generalization where static safety alignments fail to transfer to dynamic video contexts. We introduce SafeVid, a framework designed to instill video-specific safety principles in VLMMs. SafeVid uniquely transfers robust textual safety alignment capabilities to the video domain by employing detailed textual video descriptions as an interpretive bridge, facilitating LLM-based rule-driven safety reasoning. This is achieved through a closed-loop system comprising: 1) generation of SafeVid-350K, a novel 350,000-pair video-specific safety preference dataset; 2) targeted alignment of VLMMs using Direct Preference Optimization (DPO); and 3) comprehensive evaluation via our new SafeVidBench benchmark. Alignment with SafeVid-350K significantly enhances VLMM safety, with models like LLaVA-NeXT-Video demonstrating substantial improvements (e.g., up to 42.39%) on SafeVidBench. SafeVid provides critical resources and a structured approach, demonstrating that leveraging textual descriptions as a conduit for safety reasoning markedly improves the safety alignment of VLMMs. We have made SafeVid-350K dataset (https://huggingface.co/datasets/yxwang/SafeVid-350K) publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11926
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeVid: Toward Safety Aligned Video Large Multimodal Models
Wang, Yixu
Song, Jiaxin
Gao, Yifeng
Wang, Xin
Yao, Yang
Teng, Yan
Ma, Xingjun
Wang, Yingchun
Jiang, Yu-Gang
Computer Vision and Pattern Recognition
Artificial Intelligence
As Video Large Multimodal Models (VLMMs) rapidly advance, their inherent complexity introduces significant safety challenges, particularly the issue of mismatched generalization where static safety alignments fail to transfer to dynamic video contexts. We introduce SafeVid, a framework designed to instill video-specific safety principles in VLMMs. SafeVid uniquely transfers robust textual safety alignment capabilities to the video domain by employing detailed textual video descriptions as an interpretive bridge, facilitating LLM-based rule-driven safety reasoning. This is achieved through a closed-loop system comprising: 1) generation of SafeVid-350K, a novel 350,000-pair video-specific safety preference dataset; 2) targeted alignment of VLMMs using Direct Preference Optimization (DPO); and 3) comprehensive evaluation via our new SafeVidBench benchmark. Alignment with SafeVid-350K significantly enhances VLMM safety, with models like LLaVA-NeXT-Video demonstrating substantial improvements (e.g., up to 42.39%) on SafeVidBench. SafeVid provides critical resources and a structured approach, demonstrating that leveraging textual descriptions as a conduit for safety reasoning markedly improves the safety alignment of VLMMs. We have made SafeVid-350K dataset (https://huggingface.co/datasets/yxwang/SafeVid-350K) publicly available.
title SafeVid: Toward Safety Aligned Video Large Multimodal Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.11926