Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Youze, Chen, Zijun, Chen, Ruoyu, Gu, Shishen, Hu, Wenbo, Liu, Jiayang, Dong, Yinpeng, Su, Hang, Zhu, Jun, Wang, Meng, Hong, Richang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909925155799040
author Wang, Youze
Chen, Zijun
Chen, Ruoyu
Gu, Shishen
Hu, Wenbo
Liu, Jiayang
Dong, Yinpeng
Su, Hang
Zhu, Jun
Wang, Meng
Hong, Richang
author_facet Wang, Youze
Chen, Zijun
Chen, Ruoyu
Gu, Shishen
Hu, Wenbo
Liu, Jiayang
Dong, Yinpeng
Su, Hang
Zhu, Jun
Wang, Meng
Hong, Richang
contents Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability. This study introduces Trust-videoLLMs, a first comprehensive benchmark evaluating 23 state-of-the-art videoLLMs (5 commercial, 18 open-source) across five critical dimensions: truthfulness, robustness, safety, fairness, and privacy. Comprising 30 tasks with adapted, synthetic, and annotated videos, the framework assesses spatiotemporal risks, temporal consistency and cross-modal impact. Results reveal significant limitations in dynamic scene comprehension, cross-modal perturbation resilience and real-world risk mitigation. While open-source models occasionally outperform, proprietary models generally exhibit superior credibility, though scaling does not consistently improve performance. These findings underscore the need for enhanced training datat diversity and robust multimodal alignment. Trust-videoLLMs provides a publicly available, extensible toolkit for standardized trustworthiness assessments, addressing the critical gap between accuracy-focused benchmarks and demands for robustness, safety, fairness, and privacy.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12336
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
Wang, Youze
Chen, Zijun
Chen, Ruoyu
Gu, Shishen
Hu, Wenbo
Liu, Jiayang
Dong, Yinpeng
Su, Hang
Zhu, Jun
Wang, Meng
Hong, Richang
Computer Vision and Pattern Recognition
Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability. This study introduces Trust-videoLLMs, a first comprehensive benchmark evaluating 23 state-of-the-art videoLLMs (5 commercial, 18 open-source) across five critical dimensions: truthfulness, robustness, safety, fairness, and privacy. Comprising 30 tasks with adapted, synthetic, and annotated videos, the framework assesses spatiotemporal risks, temporal consistency and cross-modal impact. Results reveal significant limitations in dynamic scene comprehension, cross-modal perturbation resilience and real-world risk mitigation. While open-source models occasionally outperform, proprietary models generally exhibit superior credibility, though scaling does not consistently improve performance. These findings underscore the need for enhanced training datat diversity and robust multimodal alignment. Trust-videoLLMs provides a publicly available, extensible toolkit for standardized trustworthiness assessments, addressing the critical gap between accuracy-focused benchmarks and demands for robustness, safety, fairness, and privacy.
title Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.12336