Towards Neuro-Symbolic Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Minkyu, Goel, Harsh, Omama, Mohammad, Yang, Yunhao, Shah, Sahil, Chinchali, Sandeep
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915046300319744
author Choi, Minkyu
Goel, Harsh
Omama, Mohammad
Yang, Yunhao
Shah, Sahil
Chinchali, Sandeep
author_facet Choi, Minkyu
Goel, Harsh
Omama, Mohammad
Yang, Yunhao
Shah, Sahil
Chinchali, Sandeep
contents The unprecedented surge in video data production in recent years necessitates efficient tools to extract meaningful frames from videos for downstream tasks. Long-term temporal reasoning is a key desideratum for frame retrieval systems. While state-of-the-art foundation models, like VideoLLaMA and ViCLIP, are proficient in short-term semantic understanding, they surprisingly fail at long-term reasoning across frames. A key reason for this failure is that they intertwine per-frame perception and temporal reasoning into a single deep network. Hence, decoupling but co-designing semantic understanding and temporal reasoning is essential for efficient scene identification. We propose a system that leverages vision-language models for semantic understanding of individual frames but effectively reasons about the long-term evolution of events using state machines and temporal logic (TL) formulae that inherently capture memory. Our TL-based reasoning improves the F1 score of complex event identification by 9-15% compared to benchmarks that use GPT4 for reasoning on state-of-the-art self-driving datasets such as Waymo and NuScenes.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11021
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Neuro-Symbolic Video Understanding
Choi, Minkyu
Goel, Harsh
Omama, Mohammad
Yang, Yunhao
Shah, Sahil
Chinchali, Sandeep
Computer Vision and Pattern Recognition
Artificial Intelligence
The unprecedented surge in video data production in recent years necessitates efficient tools to extract meaningful frames from videos for downstream tasks. Long-term temporal reasoning is a key desideratum for frame retrieval systems. While state-of-the-art foundation models, like VideoLLaMA and ViCLIP, are proficient in short-term semantic understanding, they surprisingly fail at long-term reasoning across frames. A key reason for this failure is that they intertwine per-frame perception and temporal reasoning into a single deep network. Hence, decoupling but co-designing semantic understanding and temporal reasoning is essential for efficient scene identification. We propose a system that leverages vision-language models for semantic understanding of individual frames but effectively reasons about the long-term evolution of events using state machines and temporal logic (TL) formulae that inherently capture memory. Our TL-based reasoning improves the F1 score of complex event identification by 9-15% compared to benchmarks that use GPT4 for reasoning on state-of-the-art self-driving datasets such as Waymo and NuScenes.
title Towards Neuro-Symbolic Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2403.11021