Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Jiafeng, Jiang, Shixin, Dong, Xuan, Wang, Ning, Chu, Zheng, Su, Hui, Fu, Jinlan, Liu, Ming, Ng, See-Kiong, Qin, Bing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916746418454528
author Liang, Jiafeng
Jiang, Shixin
Dong, Xuan
Wang, Ning
Chu, Zheng
Su, Hui
Fu, Jinlan
Liu, Ming
Ng, See-Kiong
Qin, Bing
author_facet Liang, Jiafeng
Jiang, Shixin
Dong, Xuan
Wang, Ning
Chu, Zheng
Su, Hui
Fu, Jinlan
Liu, Ming
Ng, See-Kiong
Qin, Bing
contents Large Multimodal Models (LMMs) have recently demonstrated impressive performance on general video comprehension benchmarks. Nevertheless, for broader applications, the robustness of their temporal analysis capability needs to be thoroughly investigated yet predominantly ignored. Motivated by this, we propose a novel temporal robustness benchmark (TemRobBench), which introduces temporal inconsistency perturbations separately at the visual and textual modalities to assess the robustness of models. We evaluate 16 mainstream LMMs and find that they exhibit over-reliance on prior knowledge and textual context in adversarial environments, while ignoring the actual temporal dynamics in the video. To mitigate this issue, we design panoramic direct preference optimization (PanoDPO), which encourages LMMs to incorporate both visual and linguistic feature preferences simultaneously. Experimental results show that PanoDPO can effectively enhance the model's robustness and reliability in temporal analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency
Liang, Jiafeng
Jiang, Shixin
Dong, Xuan
Wang, Ning
Chu, Zheng
Su, Hui
Fu, Jinlan
Liu, Ming
Ng, See-Kiong
Qin, Bing
Computer Vision and Pattern Recognition
Large Multimodal Models (LMMs) have recently demonstrated impressive performance on general video comprehension benchmarks. Nevertheless, for broader applications, the robustness of their temporal analysis capability needs to be thoroughly investigated yet predominantly ignored. Motivated by this, we propose a novel temporal robustness benchmark (TemRobBench), which introduces temporal inconsistency perturbations separately at the visual and textual modalities to assess the robustness of models. We evaluate 16 mainstream LMMs and find that they exhibit over-reliance on prior knowledge and textual context in adversarial environments, while ignoring the actual temporal dynamics in the video. To mitigate this issue, we design panoramic direct preference optimization (PanoDPO), which encourages LMMs to incorporate both visual and linguistic feature preferences simultaneously. Experimental results show that PanoDPO can effectively enhance the model's robustness and reliability in temporal analysis.
title Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.14405