EvoGround: Self-Evolving Video Agents for Video Temporal Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jung, Minjoon, Zhang, Byoung-Tak, Torresani, Lorenzo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913124325523456
author Jung, Minjoon
Zhang, Byoung-Tak
Torresani, Lorenzo
author_facet Jung, Minjoon
Zhang, Byoung-Tak
Torresani, Lorenzo
contents Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the query. Existing methods rely on large, task-specific datasets requiring costly manual annotation. We introduce EvoGround, a framework of two coupled self-evolving agents, a proposer and a solver, that learn temporal grounding from raw videos without any human-labeled data. The proposer generates query--moment pairs from raw videos, while the solver learns to ground them and feeds back signals that improve the proposer in return. Through this self-reinforcing reinforcement-learning loop, the two agents are initialized from the same backbone and mutually improve across iterations. Trained on 2.5K unlabeled videos, EvoGround matches or surpasses fully supervised models across multiple VTG benchmarks, while emerging as a state-of-the-art fine-grained video captioner without manual labels.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13803
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
Jung, Minjoon
Zhang, Byoung-Tak
Torresani, Lorenzo
Computer Vision and Pattern Recognition
Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the query. Existing methods rely on large, task-specific datasets requiring costly manual annotation. We introduce EvoGround, a framework of two coupled self-evolving agents, a proposer and a solver, that learn temporal grounding from raw videos without any human-labeled data. The proposer generates query--moment pairs from raw videos, while the solver learns to ground them and feeds back signals that improve the proposer in return. Through this self-reinforcing reinforcement-learning loop, the two agents are initialized from the same backbone and mutually improve across iterations. Trained on 2.5K unlabeled videos, EvoGround matches or surpasses fully supervised models across multiple VTG benchmarks, while emerging as a state-of-the-art fine-grained video captioner without manual labels.
title EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.13803