Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruschel, Raphael, Prajapati, Hardikkumar, Rahman, Awsafur, Manjunath, B. S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909922271166464
author Ruschel, Raphael
Prajapati, Hardikkumar
Rahman, Awsafur
Manjunath, B. S.
author_facet Ruschel, Raphael
Prajapati, Hardikkumar
Rahman, Awsafur
Manjunath, B. S.
contents State-of-the-art Video Scene Graph Generation (VSGG) systems provide structured visual understanding but operate as closed, feed-forward pipelines with no ability to incorporate human guidance. In contrast, promptable segmentation models such as SAM2 enable precise user interaction but lack semantic or relational reasoning. We introduce Click2Graph, the first interactive framework for Panoptic Video Scene Graph Generation (PVSG) that unifies visual prompting with spatial, temporal, and semantic understanding. From a single user cue, such as a click or bounding box, Click2Graph segments and tracks the subject across time, autonomously discovers interacting objects, and predicts <subject, object, predicate> triplets to form a temporally consistent scene graph. Our framework introduces two key components: a Dynamic Interaction Discovery Module that generates subject-conditioned object prompts, and a Semantic Classification Head that performs joint entity and predicate reasoning. Experiments on the OpenPVSG benchmark demonstrate that Click2Graph establishes a strong foundation for user-guided PVSG, showing how human prompting can be combined with panoptic grounding and relational inference to enable controllable and interpretable video scene understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15948
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click
Ruschel, Raphael
Prajapati, Hardikkumar
Rahman, Awsafur
Manjunath, B. S.
Computer Vision and Pattern Recognition
State-of-the-art Video Scene Graph Generation (VSGG) systems provide structured visual understanding but operate as closed, feed-forward pipelines with no ability to incorporate human guidance. In contrast, promptable segmentation models such as SAM2 enable precise user interaction but lack semantic or relational reasoning. We introduce Click2Graph, the first interactive framework for Panoptic Video Scene Graph Generation (PVSG) that unifies visual prompting with spatial, temporal, and semantic understanding. From a single user cue, such as a click or bounding box, Click2Graph segments and tracks the subject across time, autonomously discovers interacting objects, and predicts <subject, object, predicate> triplets to form a temporally consistent scene graph. Our framework introduces two key components: a Dynamic Interaction Discovery Module that generates subject-conditioned object prompts, and a Semantic Classification Head that performs joint entity and predicate reasoning. Experiments on the OpenPVSG benchmark demonstrate that Click2Graph establishes a strong foundation for user-guided PVSG, showing how human prompting can be combined with panoptic grounding and relational inference to enable controllable and interpretable video scene understanding.
title Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.15948