TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Pengyu, Gorugantu, Akhil, Bhosale, Mahesh, Wasi, Abdul, Trivedi, Vishvesh, Doermann, David
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914622690295808
author Yan, Pengyu
Gorugantu, Akhil
Bhosale, Mahesh
Wasi, Abdul
Trivedi, Vishvesh
Doermann, David
author_facet Yan, Pengyu
Gorugantu, Akhil
Bhosale, Mahesh
Wasi, Abdul
Trivedi, Vishvesh
Doermann, David
contents Multi-video event understanding demands models that can locate and attribute query-relevant evidence scattered across long, heterogeneous video corpora. Existing large vision-language models (LVLMs) often underperform in this regime because they quickly exhaust their context budget and struggle to precisely localize evidentially important segments, frequently missing dense informational cues such as broadcast graphics, subtitles, and scoreboards. We introduce TRACE, an evidence grounding-guided framework that follows a ground-before-reasoning strategy for multi-video event reasoning. Our approach first builds a structured, text-searchable timeline for each video using OCR and object detection. A text-only LLM then conducts query-aware evidence localization, selecting relevant moments prior to any downstream visual reasoning. The retrieved frames and their grounding summaries are subsequently used to steer LVLM-based claim generation and cross-video citation consolidation. Experiments on MAGMaR 2026 and WikiVideo demonstrate that structured grounding markedly boosts factual completeness and attribution fidelity. On the MAGMaR validation split, TRACE raises macro-average MiRAGE F1 from 0.705 to 0.811 compared to an unguided Qwen3-VL-30B baseline, with especially strong improvements in citation recall from 0.440 to 0.628. The method also attains state-of-the-art results on the official MAGMaR 2026 leaderboard. Code is released at https://github.com/pengyu965/TRACE.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16740
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation
Yan, Pengyu
Gorugantu, Akhil
Bhosale, Mahesh
Wasi, Abdul
Trivedi, Vishvesh
Doermann, David
Computer Vision and Pattern Recognition
Multi-video event understanding demands models that can locate and attribute query-relevant evidence scattered across long, heterogeneous video corpora. Existing large vision-language models (LVLMs) often underperform in this regime because they quickly exhaust their context budget and struggle to precisely localize evidentially important segments, frequently missing dense informational cues such as broadcast graphics, subtitles, and scoreboards. We introduce TRACE, an evidence grounding-guided framework that follows a ground-before-reasoning strategy for multi-video event reasoning. Our approach first builds a structured, text-searchable timeline for each video using OCR and object detection. A text-only LLM then conducts query-aware evidence localization, selecting relevant moments prior to any downstream visual reasoning. The retrieved frames and their grounding summaries are subsequently used to steer LVLM-based claim generation and cross-video citation consolidation. Experiments on MAGMaR 2026 and WikiVideo demonstrate that structured grounding markedly boosts factual completeness and attribution fidelity. On the MAGMaR validation split, TRACE raises macro-average MiRAGE F1 from 0.705 to 0.811 compared to an unguided Qwen3-VL-30B baseline, with especially strong improvements in citation recall from 0.440 to 0.628. The method also attains state-of-the-art results on the official MAGMaR 2026 leaderboard. Code is released at https://github.com/pengyu965/TRACE.
title TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.16740