TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Pengyu, Gorugantu, Akhil, Bhosale, Mahesh, Wasi, Abdul, Trivedi, Vishvesh, Doermann, David |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering
by: Bhosale, Mahesh, et al.
Published: (2026)
by: Bhosale, Mahesh, et al.
Published: (2026)
ChartReformer: Natural Language-Driven Chart Image Editing
by: Yan, Pengyu, et al.
Published: (2024)
by: Yan, Pengyu, et al.
Published: (2024)
Score-Control for Hallucination Reduction in Diffusion Models
by: Bhosale, Mahesh, et al.
Published: (2026)
by: Bhosale, Mahesh, et al.
Published: (2026)
FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants
by: Bhosale, Mahesh, et al.
Published: (2026)
by: Bhosale, Mahesh, et al.
Published: (2026)
PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions
by: Bhosale, Mahesh, et al.
Published: (2025)
by: Bhosale, Mahesh, et al.
Published: (2025)
AutoEdit: Automatic Hyperparameter Tuning for Image Editing
by: Pham, Chau, et al.
Published: (2025)
by: Pham, Chau, et al.
Published: (2025)
TRACE: Temporal Grounding Video LLM via Causal Event Modeling
by: Guo, Yongxin, et al.
Published: (2024)
by: Guo, Yongxin, et al.
Published: (2024)
Artemis: Towards Referential Understanding in Complex Videos
by: Qiu, Jihao, et al.
Published: (2024)
by: Qiu, Jihao, et al.
Published: (2024)
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
by: Huang, Yanxiang, et al.
Published: (2026)
by: Huang, Yanxiang, et al.
Published: (2026)
TRACE: Temporal Radiology with Anatomical Change Explanation for Grounded X-ray Report Generation
by: Aranya, OFM Riaz Rahman, et al.
Published: (2026)
by: Aranya, OFM Riaz Rahman, et al.
Published: (2026)
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
by: Maniyar, Suyash, et al.
Published: (2025)
by: Maniyar, Suyash, et al.
Published: (2025)
A Survey of Video Datasets for Grounded Event Understanding
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
by: Xu, Wenhao, et al.
Published: (2025)
by: Xu, Wenhao, et al.
Published: (2025)
Context-Guided Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
by: Wu, Zhixuan, et al.
Published: (2026)
by: Wu, Zhixuan, et al.
Published: (2026)
Multi-sentence Video Grounding for Long Video Generation
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
YOLOv12: Attention-Centric Real-Time Object Detectors
by: Tian, Yunjie, et al.
Published: (2025)
by: Tian, Yunjie, et al.
Published: (2025)
TRACE: Object Motion Editing in Videos with First-Frame Trajectory Guidance
by: Phung, Quynh, et al.
Published: (2026)
by: Phung, Quynh, et al.
Published: (2026)
Personalized Large Vision-Language Models
by: Pham, Chau, et al.
Published: (2024)
by: Pham, Chau, et al.
Published: (2024)
EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation
by: Li, Bingxuan, et al.
Published: (2025)
by: Li, Bingxuan, et al.
Published: (2025)
IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
by: Zhai, Yuanhao, et al.
Published: (2024)
by: Zhai, Yuanhao, et al.
Published: (2024)
Cambrian-P: Pose-Grounded Video Understanding
by: Yang, Jihan, et al.
Published: (2026)
by: Yang, Jihan, et al.
Published: (2026)
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
by: Wang, Mengyue, et al.
Published: (2025)
by: Wang, Mengyue, et al.
Published: (2025)
Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
by: Sharma, Shivam, et al.
Published: (2026)
by: Sharma, Shivam, et al.
Published: (2026)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
by: Xie, Ming, et al.
Published: (2026)
by: Xie, Ming, et al.
Published: (2026)
Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras
by: Kong, Lingdong, et al.
Published: (2025)
by: Kong, Lingdong, et al.
Published: (2025)
Mind the Time: Temporally-Controlled Multi-Event Video Generation
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
Leaf-Based Plant Disease Detection and Explainable AI
by: Sagar, Saurav, et al.
Published: (2023)
by: Sagar, Saurav, et al.
Published: (2023)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
by: Zeng, Xiangyu, et al.
Published: (2025)
by: Zeng, Xiangyu, et al.
Published: (2025)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
Harnessing Object Grounding for Time-Sensitive Video Understanding
by: Wu, Tz-Ying, et al.
Published: (2025)
by: Wu, Tz-Ying, et al.
Published: (2025)
An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval
by: Kandhare, Mahesh, et al.
Published: (2024)
by: Kandhare, Mahesh, et al.
Published: (2024)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Enhanced Transformer-Based Tracking for Skiing Events: Overcoming Multi-Camera Challenges, Scale Variations and Rapid Motion -- SkiTB Visual Tracking Challenge 2025
by: Penta, Akhil, et al.
Published: (2025)
by: Penta, Akhil, et al.
Published: (2025)
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
by: Luo, Fuwen, et al.
Published: (2025)
by: Luo, Fuwen, et al.
Published: (2025)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
by: Xiong, Junyu, et al.
Published: (2025)
by: Xiong, Junyu, et al.
Published: (2025)
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
by: Yu, Fangxu, et al.
Published: (2026)
by: Yu, Fangxu, et al.
Published: (2026)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
When and Where do Events Switch in Multi-Event Video Generation?
by: Liao, Ruotong, et al.
Published: (2025)
by: Liao, Ruotong, et al.
Published: (2025)
Similar Items
-
CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering
by: Bhosale, Mahesh, et al.
Published: (2026) -
ChartReformer: Natural Language-Driven Chart Image Editing
by: Yan, Pengyu, et al.
Published: (2024) -
Score-Control for Hallucination Reduction in Diffusion Models
by: Bhosale, Mahesh, et al.
Published: (2026) -
FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants
by: Bhosale, Mahesh, et al.
Published: (2026) -
PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions
by: Bhosale, Mahesh, et al.
Published: (2025)