How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Shengji, Zou, Yuanhao, Zhu, Victor, Ji, Zhengping, Chen, Chen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914463194546176
author Jin, Shengji
Zou, Yuanhao
Zhu, Victor
Ji, Zhengping
Chen, Chen
author_facet Jin, Shengji
Zou, Yuanhao
Zhu, Victor
Ji, Zhengping
Chen, Chen
contents While Multimodal Large Language Models (MLLMs) have advanced Video Temporal Grounding (VTG), existing methods often couple output paradigms with different backbones, datasets, and training protocols. This makes it challenging to isolate the specific impact of the output design. Additionally, as VTG systems are increasingly considered for resource-constrained edge deployment, the trade-off between output formulation and system-level efficiency requires systematic investigation. In this paper, we present a controlled empirical study comparing three dominant VTG output paradigms: Text Numeral Generation, Temporal Token Generation, and Continuous Temporal Decoding. We evaluate these paradigms across identical compact VLMs (SmolVLM2, FastVLM, and Molmo2) using consistent datasets and LoRA fine-tuning protocols. Evaluations on Charades-STA, QVHighlights, and YouCook2 measure both localization accuracy and system efficiency, including inference latency, training throughput, and parameter overhead. Our results demonstrate that the choice of output formulation significantly affects both grounding accuracy and computational cost, independent of model scale. Specifically, the continuous distribution paradigm consistently achieves the most favorable efficiency-accuracy trade-off on the Pareto frontier, delivering robust localization with minimal latency overhead. These findings provide objective empirical guidelines for designing efficient, deployment-ready VTG systems.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08966
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms
Jin, Shengji
Zou, Yuanhao
Zhu, Victor
Ji, Zhengping
Chen, Chen
Computer Vision and Pattern Recognition
While Multimodal Large Language Models (MLLMs) have advanced Video Temporal Grounding (VTG), existing methods often couple output paradigms with different backbones, datasets, and training protocols. This makes it challenging to isolate the specific impact of the output design. Additionally, as VTG systems are increasingly considered for resource-constrained edge deployment, the trade-off between output formulation and system-level efficiency requires systematic investigation. In this paper, we present a controlled empirical study comparing three dominant VTG output paradigms: Text Numeral Generation, Temporal Token Generation, and Continuous Temporal Decoding. We evaluate these paradigms across identical compact VLMs (SmolVLM2, FastVLM, and Molmo2) using consistent datasets and LoRA fine-tuning protocols. Evaluations on Charades-STA, QVHighlights, and YouCook2 measure both localization accuracy and system efficiency, including inference latency, training throughput, and parameter overhead. Our results demonstrate that the choice of output formulation significantly affects both grounding accuracy and computational cost, independent of model scale. Specifically, the continuous distribution paradigm consistently achieves the most favorable efficiency-accuracy trade-off on the Pareto frontier, delivering robust localization with minimal latency overhead. These findings provide objective empirical guidelines for designing efficient, deployment-ready VTG systems.
title How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.08966