NeMo: Needle in a Montage for Video-Language Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Zi-Yuan, Liang, Shuo, Zheng, Duo, Li, Yanyang, Tao, Yeyao, Huang, Shijia, Feng, Wei, Qin, Jia, Yu, Jianguang, Huang, Jing, Fang, Meng, Li, Yin, Wang, Liwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908589541556224
author Hu, Zi-Yuan
Liang, Shuo
Zheng, Duo
Li, Yanyang
Tao, Yeyao
Huang, Shijia
Feng, Wei
Qin, Jia
Yu, Jianguang
Huang, Jing
Fang, Meng
Li, Yin
Wang, Liwei
author_facet Hu, Zi-Yuan
Liang, Shuo
Zheng, Duo
Li, Yanyang
Tao, Yeyao
Huang, Shijia
Feng, Wei
Qin, Jia
Yu, Jianguang
Huang, Jing
Fang, Meng
Li, Yin
Wang, Liwei
contents Recent advances in video large language models (VideoLLMs) call for new evaluation protocols and benchmarks for complex temporal reasoning in video-language understanding. Inspired by the needle in a haystack test widely used by LLMs, we introduce a novel task of Needle in a Montage (NeMo), designed to assess VideoLLMs' critical reasoning capabilities, including long-context recall and temporal grounding. To generate video question answering data for our task, we develop a scalable automated data generation pipeline that facilitates high-quality data synthesis. Built upon the proposed pipeline, we present NeMoBench, a video-language benchmark centered on our task. Specifically, our full set of NeMoBench features 31,378 automatically generated question-answer (QA) pairs from 13,486 videos with various durations ranging from seconds to hours. Experiments demonstrate that our pipeline can reliably and automatically generate high-quality evaluation data, enabling NeMoBench to be continuously updated with the latest videos. We evaluate 20 state-of-the-art models on our benchmark, providing extensive results and key insights into their capabilities and limitations. Our project page is available at: https://lavi-lab.github.io/NeMoBench.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24563
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NeMo: Needle in a Montage for Video-Language Understanding
Hu, Zi-Yuan
Liang, Shuo
Zheng, Duo
Li, Yanyang
Tao, Yeyao
Huang, Shijia
Feng, Wei
Qin, Jia
Yu, Jianguang
Huang, Jing
Fang, Meng
Li, Yin
Wang, Liwei
Computer Vision and Pattern Recognition
Computation and Language
Recent advances in video large language models (VideoLLMs) call for new evaluation protocols and benchmarks for complex temporal reasoning in video-language understanding. Inspired by the needle in a haystack test widely used by LLMs, we introduce a novel task of Needle in a Montage (NeMo), designed to assess VideoLLMs' critical reasoning capabilities, including long-context recall and temporal grounding. To generate video question answering data for our task, we develop a scalable automated data generation pipeline that facilitates high-quality data synthesis. Built upon the proposed pipeline, we present NeMoBench, a video-language benchmark centered on our task. Specifically, our full set of NeMoBench features 31,378 automatically generated question-answer (QA) pairs from 13,486 videos with various durations ranging from seconds to hours. Experiments demonstrate that our pipeline can reliably and automatically generate high-quality evaluation data, enabling NeMoBench to be continuously updated with the latest videos. We evaluate 20 state-of-the-art models on our benchmark, providing extensive results and key insights into their capabilities and limitations. Our project page is available at: https://lavi-lab.github.io/NeMoBench.
title NeMo: Needle in a Montage for Video-Language Understanding
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2509.24563