LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jingyi, Ju, Jianzhong, Luan, Jian, Deng, Zhidong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914929385144320
author Wang, Jingyi
Ju, Jianzhong
Luan, Jian
Deng, Zhidong
author_facet Wang, Jingyi
Ju, Jianzhong
Luan, Jian
Deng, Zhidong
contents Recent advances in large vision-language models (VLMs) typically employ vision encoders based on the Vision Transformer (ViT) architecture. The division of the images into patches by ViT results in a fragmented perception, thereby hindering the visual understanding capabilities of VLMs. In this paper, we propose an innovative enhancement to address this limitation by introducing a Scene Graph Expression (SGE) module in VLMs. This module extracts and structurally expresses the complex semantic information within images, thereby improving the foundational perception and understanding abilities of VLMs. Extensive experiments demonstrate that integrating our SGE module significantly enhances the VLM's performance in vision-language tasks, indicating its effectiveness in preserving intricate semantic details and facilitating better visual understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16224
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
Wang, Jingyi
Ju, Jianzhong
Luan, Jian
Deng, Zhidong
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in large vision-language models (VLMs) typically employ vision encoders based on the Vision Transformer (ViT) architecture. The division of the images into patches by ViT results in a fragmented perception, thereby hindering the visual understanding capabilities of VLMs. In this paper, we propose an innovative enhancement to address this limitation by introducing a Scene Graph Expression (SGE) module in VLMs. This module extracts and structurally expresses the complex semantic information within images, thereby improving the foundational perception and understanding abilities of VLMs. Extensive experiments demonstrate that integrating our SGE module significantly enhances the VLM's performance in vision-language tasks, indicating its effectiveness in preserving intricate semantic details and facilitating better visual understanding.
title LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2408.16224