Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Youze, Chen, Zijun, Chen, Ruoyu, Gu, Shishen, Hu, Wenbo, Liu, Jiayang, Dong, Yinpeng, Su, Hang, Zhu, Jun, Wang, Meng, Hong, Richang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-training Models with Contrastive Learning
by: Wang, Youze, et al.
Published: (2023)
by: Wang, Youze, et al.
Published: (2023)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
Iterative Adversarial Attack on Image-guided Story Ending Generation
by: Wang, Youze, et al.
Published: (2023)
by: Wang, Youze, et al.
Published: (2023)
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
by: Yang, Xiao, et al.
Published: (2025)
by: Yang, Xiao, et al.
Published: (2025)
Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning
by: Chen, Zijun, et al.
Published: (2025)
by: Chen, Zijun, et al.
Published: (2025)
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models
by: Chen, Zijun, et al.
Published: (2024)
by: Chen, Zijun, et al.
Published: (2024)
Towards the Worst-case Robustness of Large Language Models
by: Chen, Huanran, et al.
Published: (2025)
by: Chen, Huanran, et al.
Published: (2025)
Exploring the Transferability of Visual Prompting for Multimodal Large Language Models
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
Rethinking Model Ensemble in Transfer-based Adversarial Attacks
by: Chen, Huanran, et al.
Published: (2023)
by: Chen, Huanran, et al.
Published: (2023)
Robust Classification via a Single Diffusion Model
by: Chen, Huanran, et al.
Published: (2023)
by: Chen, Huanran, et al.
Published: (2023)
From Static to Dynamic: Adapting Landmark-Aware Image Models for Facial Expression Recognition in Videos
by: Chen, Yin, et al.
Published: (2023)
by: Chen, Yin, et al.
Published: (2023)
RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations
by: Li, Hanyu, et al.
Published: (2026)
by: Li, Hanyu, et al.
Published: (2026)
TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models
by: Sun, Ruoyu, et al.
Published: (2025)
by: Sun, Ruoyu, et al.
Published: (2025)
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Your Diffusion Model is Secretly a Certifiably Robust Classifier
by: Chen, Huanran, et al.
Published: (2024)
by: Chen, Huanran, et al.
Published: (2024)
Unveiling the Basin-Like Loss Landscape in Large Language Models
by: Chen, Huanran, et al.
Published: (2025)
by: Chen, Huanran, et al.
Published: (2025)
Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data
by: Chen, Yin, et al.
Published: (2024)
by: Chen, Yin, et al.
Published: (2024)
Deep sub-ensembles meets quantile regression: uncertainty-aware imputation for time series
by: Liu, Ying, et al.
Published: (2023)
by: Liu, Ying, et al.
Published: (2023)
Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
FaceCat: Enhancing Face Recognition Security with a Unified Diffusion Model
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Embodied Active Defense: Leveraging Recurrent Feedback to Counter Adversarial Patches
by: Wu, Lingxuan, et al.
Published: (2024)
by: Wu, Lingxuan, et al.
Published: (2024)
BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators
by: Tian, Yu, et al.
Published: (2024)
by: Tian, Yu, et al.
Published: (2024)
Correspondence on “Needs Assessment for Updating Institute of Medicine Standards for Trustworthy Clinical Practice Guidelines”
by: Zijun Wang, et al.
Published: (2025)
by: Zijun Wang, et al.
Published: (2025)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
by: Li, Huiqiong, et al.
Published: (2026)
by: Li, Huiqiong, et al.
Published: (2026)
Exploring the Generalizability of Factual Hallucination Mitigation via Enhancing Precise Knowledge Utilization
by: Zhang, Siyuan, et al.
Published: (2025)
by: Zhang, Siyuan, et al.
Published: (2025)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
by: Feng, Xiang, et al.
Published: (2026)
by: Feng, Xiang, et al.
Published: (2026)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
by: Li, Keliang, et al.
Published: (2024)
by: Li, Keliang, et al.
Published: (2024)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
by: Pan, Yaning, et al.
Published: (2025)
by: Pan, Yaning, et al.
Published: (2025)
A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
by: Su, Hang, et al.
Published: (2025)
by: Su, Hang, et al.
Published: (2025)
Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context Learning
by: Huang, Zhuo, et al.
Published: (2023)
by: Huang, Zhuo, et al.
Published: (2023)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
by: Chen, Houlun, et al.
Published: (2024)
by: Chen, Houlun, et al.
Published: (2024)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding
by: Wei, Wenbo, et al.
Published: (2024)
by: Wei, Wenbo, et al.
Published: (2024)
Structured Attention Matters to Multimodal LLMs in Document Understanding
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
A dilated convolution‐based method with time series fine tuning for data‐driven crack length estimation
by: Jiaxin Gao, et al.
Published: (2024)
by: Jiaxin Gao, et al.
Published: (2024)
Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection
by: Yakun, Cui, et al.
Published: (2025)
by: Yakun, Cui, et al.
Published: (2025)
DIFFender: Diffusion-Based Adversarial Defense against Patch Attacks
by: Kang, Caixin, et al.
Published: (2023)
by: Kang, Caixin, et al.
Published: (2023)
Similar Items
-
Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-training Models with Contrastive Learning
by: Wang, Youze, et al.
Published: (2023) -
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
by: Wang, Youze, et al.
Published: (2025) -
Iterative Adversarial Attack on Image-guided Story Ending Generation
by: Wang, Youze, et al.
Published: (2023) -
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
by: Yang, Xiao, et al.
Published: (2025) -
Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning
by: Chen, Zijun, et al.
Published: (2025)