Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhou, Lin, Joe, Aakur\\, Sathyanarayanan N.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917124081975296
author Chen, Zhou
Lin, Joe
Aakur\\, Sathyanarayanan N.
author_facet Chen, Zhou
Lin, Joe
Aakur\\, Sathyanarayanan N.
contents Humans naturally perceive continuous experience as a hierarchy of temporally nested events, fine-grained actions embedded within coarser routines. Replicating this structure in computer vision requires models that can segment video not just retrospectively, but predictively and hierarchically. We introduce PARSE, a unified framework that learns multiscale event structure directly from streaming video without supervision. PARSE organizes perception into a hierarchy of recurrent predictors, each operating at its own temporal granularity: lower layers model short-term dynamics while higher layers integrate longer-term context through attention-based feedback. Event boundaries emerge naturally as transient peaks in prediction error, yielding temporally coherent, nested partonomies that mirror the containment relations observed in human event perception. Evaluated across three benchmarks, Breakfast Actions, 50 Salads, and Assembly 101, PARSE achieves state-of-the-art performance among streaming methods and rivals offline baselines in both temporal alignment (H-GEBD) and structural consistency (TED, hF1). The results demonstrate that predictive learning under uncertainty provides a scalable path toward human-like temporal abstraction and compositional event understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04219
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
Chen, Zhou
Lin, Joe
Aakur\\, Sathyanarayanan N.
Computer Vision and Pattern Recognition
Humans naturally perceive continuous experience as a hierarchy of temporally nested events, fine-grained actions embedded within coarser routines. Replicating this structure in computer vision requires models that can segment video not just retrospectively, but predictively and hierarchically. We introduce PARSE, a unified framework that learns multiscale event structure directly from streaming video without supervision. PARSE organizes perception into a hierarchy of recurrent predictors, each operating at its own temporal granularity: lower layers model short-term dynamics while higher layers integrate longer-term context through attention-based feedback. Event boundaries emerge naturally as transient peaks in prediction error, yielding temporally coherent, nested partonomies that mirror the containment relations observed in human event perception. Evaluated across three benchmarks, Breakfast Actions, 50 Salads, and Assembly 101, PARSE achieves state-of-the-art performance among streaming methods and rivals offline baselines in both temporal alignment (H-GEBD) and structural consistency (TED, hF1). The results demonstrate that predictive learning under uncertainty provides a scalable path toward human-like temporal abstraction and compositional event understanding.
title Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.04219