NeMo: Needle in a Montage for Video-Language Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908589541556224 |
|---|---|
| author | Hu, Zi-Yuan Liang, Shuo Zheng, Duo Li, Yanyang Tao, Yeyao Huang, Shijia Feng, Wei Qin, Jia Yu, Jianguang Huang, Jing Fang, Meng Li, Yin Wang, Liwei |
| author_facet | Hu, Zi-Yuan Liang, Shuo Zheng, Duo Li, Yanyang Tao, Yeyao Huang, Shijia Feng, Wei Qin, Jia Yu, Jianguang Huang, Jing Fang, Meng Li, Yin Wang, Liwei |
| contents | Recent advances in video large language models (VideoLLMs) call for new evaluation protocols and benchmarks for complex temporal reasoning in video-language understanding. Inspired by the needle in a haystack test widely used by LLMs, we introduce a novel task of Needle in a Montage (NeMo), designed to assess VideoLLMs' critical reasoning capabilities, including long-context recall and temporal grounding. To generate video question answering data for our task, we develop a scalable automated data generation pipeline that facilitates high-quality data synthesis. Built upon the proposed pipeline, we present NeMoBench, a video-language benchmark centered on our task. Specifically, our full set of NeMoBench features 31,378 automatically generated question-answer (QA) pairs from 13,486 videos with various durations ranging from seconds to hours. Experiments demonstrate that our pipeline can reliably and automatically generate high-quality evaluation data, enabling NeMoBench to be continuously updated with the latest videos. We evaluate 20 state-of-the-art models on our benchmark, providing extensive results and key insights into their capabilities and limitations. Our project page is available at: https://lavi-lab.github.io/NeMoBench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_24563 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | NeMo: Needle in a Montage for Video-Language Understanding Hu, Zi-Yuan Liang, Shuo Zheng, Duo Li, Yanyang Tao, Yeyao Huang, Shijia Feng, Wei Qin, Jia Yu, Jianguang Huang, Jing Fang, Meng Li, Yin Wang, Liwei Computer Vision and Pattern Recognition Computation and Language Recent advances in video large language models (VideoLLMs) call for new evaluation protocols and benchmarks for complex temporal reasoning in video-language understanding. Inspired by the needle in a haystack test widely used by LLMs, we introduce a novel task of Needle in a Montage (NeMo), designed to assess VideoLLMs' critical reasoning capabilities, including long-context recall and temporal grounding. To generate video question answering data for our task, we develop a scalable automated data generation pipeline that facilitates high-quality data synthesis. Built upon the proposed pipeline, we present NeMoBench, a video-language benchmark centered on our task. Specifically, our full set of NeMoBench features 31,378 automatically generated question-answer (QA) pairs from 13,486 videos with various durations ranging from seconds to hours. Experiments demonstrate that our pipeline can reliably and automatically generate high-quality evaluation data, enabling NeMoBench to be continuously updated with the latest videos. We evaluate 20 state-of-the-art models on our benchmark, providing extensive results and key insights into their capabilities and limitations. Our project page is available at: https://lavi-lab.github.io/NeMoBench. |
| title | NeMo: Needle in a Montage for Video-Language Understanding |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2509.24563 |