Towards Lifelong Scene Graph Generation with Knowledge-ware In-context Prompt Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Tao, Wu, Tongtong, Zhang, Dongyang, Duan, Guiduo, Qin, Ke, Li, Yuan-Fang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911764932722688
author He, Tao
Wu, Tongtong
Zhang, Dongyang
Duan, Guiduo
Qin, Ke
Li, Yuan-Fang
author_facet He, Tao
Wu, Tongtong
Zhang, Dongyang
Duan, Guiduo
Qin, Ke
Li, Yuan-Fang
contents Scene graph generation (SGG) endeavors to predict visual relationships between pairs of objects within an image. Prevailing SGG methods traditionally assume a one-off learning process for SGG. This conventional paradigm may necessitate repetitive training on all previously observed samples whenever new relationships emerge, mitigating the risk of forgetting previously acquired knowledge. This work seeks to address this pitfall inherent in a suite of prior relationship predictions. Motivated by the achievements of in-context learning in pretrained language models, our approach imbues the model with the capability to predict relationships and continuously acquire novel knowledge without succumbing to catastrophic forgetting. To achieve this goal, we introduce a novel and pragmatic framework for scene graph generation, namely Lifelong Scene Graph Generation (LSGG), where tasks, such as predicates, unfold in a streaming fashion. In this framework, the model is constrained to exclusive training on the present task, devoid of access to previously encountered training data, except for a limited number of exemplars, but the model is tasked with inferring all predicates it has encountered thus far. Rigorous experiments demonstrate the superiority of our proposed method over state-of-the-art SGG models in the context of LSGG across a diverse array of metrics. Besides, extensive experiments on the two mainstream benchmark datasets, VG and Open-Image(v6), show the superiority of our proposed model to a number of competitive SGG models in terms of continuous learning and conventional settings. Moreover, comprehensive ablation experiments demonstrate the effectiveness of each component in our model.
format Preprint
id arxiv_https___arxiv_org_abs_2401_14626
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Lifelong Scene Graph Generation with Knowledge-ware In-context Prompt Learning
He, Tao
Wu, Tongtong
Zhang, Dongyang
Duan, Guiduo
Qin, Ke
Li, Yuan-Fang
Computer Vision and Pattern Recognition
Scene graph generation (SGG) endeavors to predict visual relationships between pairs of objects within an image. Prevailing SGG methods traditionally assume a one-off learning process for SGG. This conventional paradigm may necessitate repetitive training on all previously observed samples whenever new relationships emerge, mitigating the risk of forgetting previously acquired knowledge. This work seeks to address this pitfall inherent in a suite of prior relationship predictions. Motivated by the achievements of in-context learning in pretrained language models, our approach imbues the model with the capability to predict relationships and continuously acquire novel knowledge without succumbing to catastrophic forgetting. To achieve this goal, we introduce a novel and pragmatic framework for scene graph generation, namely Lifelong Scene Graph Generation (LSGG), where tasks, such as predicates, unfold in a streaming fashion. In this framework, the model is constrained to exclusive training on the present task, devoid of access to previously encountered training data, except for a limited number of exemplars, but the model is tasked with inferring all predicates it has encountered thus far. Rigorous experiments demonstrate the superiority of our proposed method over state-of-the-art SGG models in the context of LSGG across a diverse array of metrics. Besides, extensive experiments on the two mainstream benchmark datasets, VG and Open-Image(v6), show the superiority of our proposed model to a number of competitive SGG models in terms of continuous learning and conventional settings. Moreover, comprehensive ablation experiments demonstrate the effectiveness of each component in our model.
title Towards Lifelong Scene Graph Generation with Knowledge-ware In-context Prompt Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.14626