Learning Object State Changes in Videos: An Open-World Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xue, Zihui, Ashutosh, Kumar, Grauman, Kristen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916190606065664
author Xue, Zihui
Ashutosh, Kumar
Grauman, Kristen
author_facet Xue, Zihui
Ashutosh, Kumar
Grauman, Kristen
contents Object State Changes (OSCs) are pivotal for video understanding. While humans can effortlessly generalize OSC understanding from familiar to unknown objects, current approaches are confined to a closed vocabulary. Addressing this gap, we introduce a novel open-world formulation for the video OSC problem. The goal is to temporally localize the three stages of an OSC -- the object's initial state, its transitioning state, and its end state -- whether or not the object has been observed during training. Towards this end, we develop VidOSC, a holistic learning approach that: (1) leverages text and vision-language models for supervisory signals to obviate manually labeling OSC training data, and (2) abstracts fine-grained shared state representations from objects to enhance generalization. Furthermore, we present HowToChange, the first open-world benchmark for video OSC localization, which offers an order of magnitude increase in the label space and annotation volume compared to the best existing benchmark. Experimental results demonstrate the efficacy of our approach, in both traditional closed-world and open-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2312_11782
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning Object State Changes in Videos: An Open-World Perspective
Xue, Zihui
Ashutosh, Kumar
Grauman, Kristen
Computer Vision and Pattern Recognition
Object State Changes (OSCs) are pivotal for video understanding. While humans can effortlessly generalize OSC understanding from familiar to unknown objects, current approaches are confined to a closed vocabulary. Addressing this gap, we introduce a novel open-world formulation for the video OSC problem. The goal is to temporally localize the three stages of an OSC -- the object's initial state, its transitioning state, and its end state -- whether or not the object has been observed during training. Towards this end, we develop VidOSC, a holistic learning approach that: (1) leverages text and vision-language models for supervisory signals to obviate manually labeling OSC training data, and (2) abstracts fine-grained shared state representations from objects to enhance generalization. Furthermore, we present HowToChange, the first open-world benchmark for video OSC localization, which offers an order of magnitude increase in the label space and annotation volume compared to the best existing benchmark. Experimental results demonstrate the efficacy of our approach, in both traditional closed-world and open-world scenarios.
title Learning Object State Changes in Videos: An Open-World Perspective
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.11782