LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Patel, Alkesh, Ozyildirim, Melis, Cheng, Ying-Chang, Nagarajan, Ganesh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911584978206720
author Patel, Alkesh
Ozyildirim, Melis
Cheng, Ying-Chang
Nagarajan, Ganesh
author_facet Patel, Alkesh
Ozyildirim, Melis
Cheng, Ying-Chang
Nagarajan, Ganesh
contents Long video summarization presents significant challenges for current multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded. In this work, we present LVSum, a human-annotated benchmark designed specifically for evaluating long video summarization with fine-grained temporal alignment. LVSum comprises diverse long-form videos across 13 domains, each paired with human-generated summaries containing precise temporal references. We conduct a comprehensive evaluation of both proprietary and open-source MLLMs on LVSum, assessing performance using newly introduced LLM-based metrics for content relevance and modality coherence, alongside standard evaluation metrics. Our experiments reveal systematic gaps in temporal understanding among existing MLLMs and offer insights that establish a new foundation for advancing temporal reasoning in long video summarization.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10024
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
Patel, Alkesh
Ozyildirim, Melis
Cheng, Ying-Chang
Nagarajan, Ganesh
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Long video summarization presents significant challenges for current multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded. In this work, we present LVSum, a human-annotated benchmark designed specifically for evaluating long video summarization with fine-grained temporal alignment. LVSum comprises diverse long-form videos across 13 domains, each paired with human-generated summaries containing precise temporal references. We conduct a comprehensive evaluation of both proprietary and open-source MLLMs on LVSum, assessing performance using newly introduced LLM-based metrics for content relevance and modality coherence, alongside standard evaluation metrics. Our experiments reveal systematic gaps in temporal understanding among existing MLLMs and offer insights that establish a new foundation for advancing temporal reasoning in long video summarization.
title LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2604.10024