Look, Remember and Reason: Grounded reasoning in videos with language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhattacharyya, Apratim, Panchal, Sunny, Lee, Mingu, Pourreza, Reza, Madan, Pulkit, Memisevic, Roland
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911761461936128
author Bhattacharyya, Apratim
Panchal, Sunny
Lee, Mingu
Pourreza, Reza
Madan, Pulkit
Memisevic, Roland
author_facet Bhattacharyya, Apratim
Panchal, Sunny
Lee, Mingu
Pourreza, Reza
Madan, Pulkit
Memisevic, Roland
contents Multi-modal language models (LM) have recently shown promising performance in high-level reasoning tasks on videos. However, existing methods still fall short in tasks like causal or compositional spatiotemporal reasoning over actions, in which model predictions need to be grounded in fine-grained low-level details, such as object motions and object interactions. In this work, we propose training an LM end-to-end on low-level surrogate tasks, including object detection, re-identification, and tracking, to endow the model with the required low-level visual capabilities. We show that a two-stream video encoder with spatiotemporal attention is effective at capturing the required static and motion-based cues in the video. By leveraging the LM's ability to perform the low-level surrogate tasks, we can cast reasoning in videos as the three-step process of Look, Remember, Reason wherein visual information is extracted using low-level visual skills step-by-step and then integrated to arrive at a final answer. We demonstrate the effectiveness of our framework on diverse visual reasoning tasks from the ACRE, CATER, Something-Else and STAR datasets. Our approach is trainable end-to-end and surpasses state-of-the-art task-specific methods across these tasks by a large margin.
format Preprint
id arxiv_https___arxiv_org_abs_2306_17778
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Look, Remember and Reason: Grounded reasoning in videos with language models
Bhattacharyya, Apratim
Panchal, Sunny
Lee, Mingu
Pourreza, Reza
Madan, Pulkit
Memisevic, Roland
Computer Vision and Pattern Recognition
Machine Learning
Multi-modal language models (LM) have recently shown promising performance in high-level reasoning tasks on videos. However, existing methods still fall short in tasks like causal or compositional spatiotemporal reasoning over actions, in which model predictions need to be grounded in fine-grained low-level details, such as object motions and object interactions. In this work, we propose training an LM end-to-end on low-level surrogate tasks, including object detection, re-identification, and tracking, to endow the model with the required low-level visual capabilities. We show that a two-stream video encoder with spatiotemporal attention is effective at capturing the required static and motion-based cues in the video. By leveraging the LM's ability to perform the low-level surrogate tasks, we can cast reasoning in videos as the three-step process of Look, Remember, Reason wherein visual information is extracted using low-level visual skills step-by-step and then integrated to arrive at a final answer. We demonstrate the effectiveness of our framework on diverse visual reasoning tasks from the ACRE, CATER, Something-Else and STAR datasets. Our approach is trainable end-to-end and surpasses state-of-the-art task-specific methods across these tasks by a large margin.
title Look, Remember and Reason: Grounded reasoning in videos with language models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2306.17778