What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Xuanming, Bhoi, Jaiminkumar Ashokbhai, Peng, Chionh Wei, Kuek, Adriel, Lim, Ser Nam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912284016640000
author Cui, Xuanming
Bhoi, Jaiminkumar Ashokbhai
Peng, Chionh Wei
Kuek, Adriel
Lim, Ser Nam
author_facet Cui, Xuanming
Bhoi, Jaiminkumar Ashokbhai
Peng, Chionh Wei
Kuek, Adriel
Lim, Ser Nam
contents Dynamic Scene Graph Generation (DSGG) for videos is a challenging task in computer vision. While existing approaches often focus on sophisticated architectural design and solely use recall during evaluation, we take a closer look at their predicted scene graphs and discover three critical issues with existing DSGG methods: severe precision-recall trade-off, lack of awareness on triplet importance, and inappropriate evaluation protocols. On the other hand, recent advances of Large Multimodal Models (LMMs) have shown great capabilities in video understanding, yet they have not been tested on fine-grained, frame-wise understanding tasks like DSGG. In this work, we conduct the first systematic analysis of Video LMMs for performing DSGG. Without relying on sophisticated architectural design, we show that LMMs with simple decoder-only structure can be turned into State-of-the-Art scene graph generators that effectively overcome the aforementioned issues, while requiring little finetuning (5-10% training data).
format Preprint
id arxiv_https___arxiv_org_abs_2503_15846
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
Cui, Xuanming
Bhoi, Jaiminkumar Ashokbhai
Peng, Chionh Wei
Kuek, Adriel
Lim, Ser Nam
Computer Vision and Pattern Recognition
Dynamic Scene Graph Generation (DSGG) for videos is a challenging task in computer vision. While existing approaches often focus on sophisticated architectural design and solely use recall during evaluation, we take a closer look at their predicted scene graphs and discover three critical issues with existing DSGG methods: severe precision-recall trade-off, lack of awareness on triplet importance, and inappropriate evaluation protocols. On the other hand, recent advances of Large Multimodal Models (LMMs) have shown great capabilities in video understanding, yet they have not been tested on fine-grained, frame-wise understanding tasks like DSGG. In this work, we conduct the first systematic analysis of Video LMMs for performing DSGG. Without relying on sophisticated architectural design, we show that LMMs with simple decoder-only structure can be turned into State-of-the-Art scene graph generators that effectively overcome the aforementioned issues, while requiring little finetuning (5-10% training data).
title What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.15846