SPOC: Spatially-Progressing Object State Change Segmentation in Video

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mandikal, Priyanka, Nagarajan, Tushar, Stoken, Alex, Xue, Zihui, Grauman, Kristen
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911335631028224
author Mandikal, Priyanka
Nagarajan, Tushar
Stoken, Alex
Xue, Zihui
Grauman, Kristen
author_facet Mandikal, Priyanka
Nagarajan, Tushar
Stoken, Alex
Xue, Zihui
Grauman, Kristen
contents Object state changes in video reveal critical cues about human and agent activity. However, existing methods are limited to temporal localization of when the object is in its initial state (e.g., cheese block) versus when it has completed a state change (e.g., grated cheese), offering no insight into where the change is unfolding. We propose to deepen the problem by introducing the spatially-progressing object state change segmentation task. The goal is to segment at the pixel-level those regions of an object that are actionable and those that are transformed. We show that state-of-the-art VLMs and video segmentation methods struggle at this task, underscoring its difficulty and novelty. As an initial baseline, we design a VLM-based pseudo-labeling approach, state-change dynamics constraints, and a novel WhereToChange benchmark built on in-the-wild Internet videos. Experiments on two datasets validate both the challenge of the new task as well as the promise of our model for localizing exactly where and how fast objects are changing in video. We further demonstrate useful implications for tracking activity progress to benefit robotic agents. Overall, our work positions spatial OSC segmentation as a new frontier task for video understanding: one that challenges current SOTA methods and invites the community to build more robust, state-change-sensitive representations. Project page: https://vision.cs.utexas.edu/projects/spoc-spatially-progressing-osc
format Preprint
id arxiv_https___arxiv_org_abs_2503_11953
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPOC: Spatially-Progressing Object State Change Segmentation in Video
Mandikal, Priyanka
Nagarajan, Tushar
Stoken, Alex
Xue, Zihui
Grauman, Kristen
Computer Vision and Pattern Recognition
Object state changes in video reveal critical cues about human and agent activity. However, existing methods are limited to temporal localization of when the object is in its initial state (e.g., cheese block) versus when it has completed a state change (e.g., grated cheese), offering no insight into where the change is unfolding. We propose to deepen the problem by introducing the spatially-progressing object state change segmentation task. The goal is to segment at the pixel-level those regions of an object that are actionable and those that are transformed. We show that state-of-the-art VLMs and video segmentation methods struggle at this task, underscoring its difficulty and novelty. As an initial baseline, we design a VLM-based pseudo-labeling approach, state-change dynamics constraints, and a novel WhereToChange benchmark built on in-the-wild Internet videos. Experiments on two datasets validate both the challenge of the new task as well as the promise of our model for localizing exactly where and how fast objects are changing in video. We further demonstrate useful implications for tracking activity progress to benefit robotic agents. Overall, our work positions spatial OSC segmentation as a new frontier task for video understanding: one that challenges current SOTA methods and invites the community to build more robust, state-change-sensitive representations. Project page: https://vision.cs.utexas.edu/projects/spoc-spatially-progressing-osc
title SPOC: Spatially-Progressing Object State Change Segmentation in Video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.11953