Unhackable Temporal Rewarding for Scalable Video MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, En, Lin, Kangheng, Zhao, Liang, Wei, Yana, Zhu, Zining, Wei, Haoran, Sun, Jianjian, Ge, Zheng, Zhang, Xiangyu, Wang, Jingyu, Tao, Wenbing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909497749929984
author Yu, En
Lin, Kangheng
Zhao, Liang
Wei, Yana
Zhu, Zining
Wei, Haoran
Sun, Jianjian
Ge, Zheng
Zhang, Xiangyu
Wang, Jingyu
Tao, Wenbing
author_facet Yu, En
Lin, Kangheng
Zhao, Liang
Wei, Yana
Zhu, Zining
Wei, Haoran
Sun, Jianjian
Ge, Zheng
Zhang, Xiangyu
Wang, Jingyu
Tao, Wenbing
contents In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the "anti-scaling law", where more data and larger models lead to worse performance. This study unmasks the culprit: "temporal hacking", a phenomenon where models shortcut by fixating on select frames, missing the full video narrative. In this work, we systematically establish a comprehensive theory of temporal hacking, defining it from a reinforcement learning perspective, introducing the Temporal Perplexity (TPL) score to assess this misalignment, and proposing the Unhackable Temporal Rewarding (UTR) framework to mitigate the temporal hacking. Both theoretically and empirically, TPL proves to be a reliable indicator of temporal modeling quality, correlating strongly with frame activation patterns. Extensive experiments reveal that UTR not only counters temporal hacking but significantly elevates video comprehension capabilities. This work not only advances video-AI systems but also illuminates the critical importance of aligning proxy rewards with true objectives in MLLM development.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12081
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unhackable Temporal Rewarding for Scalable Video MLLMs
Yu, En
Lin, Kangheng
Zhao, Liang
Wei, Yana
Zhu, Zining
Wei, Haoran
Sun, Jianjian
Ge, Zheng
Zhang, Xiangyu
Wang, Jingyu
Tao, Wenbing
Computer Vision and Pattern Recognition
Computation and Language
In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the "anti-scaling law", where more data and larger models lead to worse performance. This study unmasks the culprit: "temporal hacking", a phenomenon where models shortcut by fixating on select frames, missing the full video narrative. In this work, we systematically establish a comprehensive theory of temporal hacking, defining it from a reinforcement learning perspective, introducing the Temporal Perplexity (TPL) score to assess this misalignment, and proposing the Unhackable Temporal Rewarding (UTR) framework to mitigate the temporal hacking. Both theoretically and empirically, TPL proves to be a reliable indicator of temporal modeling quality, correlating strongly with frame activation patterns. Extensive experiments reveal that UTR not only counters temporal hacking but significantly elevates video comprehension capabilities. This work not only advances video-AI systems but also illuminates the critical importance of aligning proxy rewards with true objectives in MLLM development.
title Unhackable Temporal Rewarding for Scalable Video MLLMs
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2502.12081