Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Minseok, Lee, Minhyeok, Kim, Minjung, Lee, Jungho, Kim, Donghyeong, Woo, Sungmin, Jeon, Inseok, Lee, Sangyoun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917357708902400
author Kang, Minseok
Lee, Minhyeok
Kim, Minjung
Lee, Jungho
Kim, Donghyeong
Woo, Sungmin
Jeon, Inseok
Lee, Sangyoun
author_facet Kang, Minseok
Lee, Minhyeok
Kim, Minjung
Lee, Jungho
Kim, Donghyeong
Woo, Sungmin
Jeon, Inseok
Lee, Sangyoun
contents Weakly-supervised video scene graph generation (WS-VSGG) aims to parse video content into structured relational triplets without bounding box annotations and with only sparse temporal labeling, significantly reducing annotation costs. Without ground-truth bounding boxes, these methods rely on off-the-shelf detectors to generate object proposals, yet largely overlook a fundamental discrepancy from fullysupervised pipelines. Fully-supervised detectors implicitly filter out noninteractive objects, while off-the-shelf detectors indiscriminately detect all visible objects, overwhelming relation models with noisy pairs.We address this by introducing a learnable pair affinity that estimates the likelihood of interaction between subject-object pairs. Through Pair Affinity Learning and Scoring (PALS), pair affinity is incorporated into inferencetime ranking and further integrated into contextual reasoning through Pair Affinity Modulation (PAM), enabling the model to suppress noninteractive pairs and focus on relationally meaningful ones. To provide cleaner supervision for pair affinity learning, we further propose Relation- Aware Matching (RAM), which leverages vision-language grounding to resolve class-level ambiguity in pseudo-label generation. Extensive experiments on Action Genome demonstrate that our approach consistently yields substantial improvements across different baselines and backbones, achieving state-of-the-art WS-VSGG performance.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21559
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning
Kang, Minseok
Lee, Minhyeok
Kim, Minjung
Lee, Jungho
Kim, Donghyeong
Woo, Sungmin
Jeon, Inseok
Lee, Sangyoun
Computer Vision and Pattern Recognition
Weakly-supervised video scene graph generation (WS-VSGG) aims to parse video content into structured relational triplets without bounding box annotations and with only sparse temporal labeling, significantly reducing annotation costs. Without ground-truth bounding boxes, these methods rely on off-the-shelf detectors to generate object proposals, yet largely overlook a fundamental discrepancy from fullysupervised pipelines. Fully-supervised detectors implicitly filter out noninteractive objects, while off-the-shelf detectors indiscriminately detect all visible objects, overwhelming relation models with noisy pairs.We address this by introducing a learnable pair affinity that estimates the likelihood of interaction between subject-object pairs. Through Pair Affinity Learning and Scoring (PALS), pair affinity is incorporated into inferencetime ranking and further integrated into contextual reasoning through Pair Affinity Modulation (PAM), enabling the model to suppress noninteractive pairs and focus on relationally meaningful ones. To provide cleaner supervision for pair affinity learning, we further propose Relation- Aware Matching (RAM), which leverages vision-language grounding to resolve class-level ambiguity in pseudo-label generation. Extensive experiments on Action Genome demonstrate that our approach consistently yields substantial improvements across different baselines and backbones, achieving state-of-the-art WS-VSGG performance.
title Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.21559