ELMM: Efficient Lightweight Multimodal Large Language Models for Multimodal Knowledge Graph Completion

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huang, Wei, Li, Peining, Liang, Meiyu, Hou, Xu, Du, Junping, Shao, Yingxia, Ye, Guanhua, Liu, Wu, Lu, Kangkang, Yu, Yang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918274009137152
author Huang, Wei
Li, Peining
Liang, Meiyu
Hou, Xu
Du, Junping
Shao, Yingxia
Ye, Guanhua
Liu, Wu
Lu, Kangkang
Yu, Yang
author_facet Huang, Wei
Li, Peining
Liang, Meiyu
Hou, Xu
Du, Junping
Shao, Yingxia
Ye, Guanhua
Liu, Wu
Lu, Kangkang
Yu, Yang
contents Multimodal Knowledge Graphs (MKGs) extend traditional knowledge graphs by incorporating visual and textual modalities, enabling richer and more expressive entity representations. However, existing MKGs often suffer from incompleteness, which hinder their effectiveness in downstream tasks. Therefore, multimodal knowledge graph completion (MKGC) task is receiving increasing attention. While large language models (LLMs) have shown promise for knowledge graph completion (KGC), their application to the multimodal setting remains underexplored. Moreover, applying Multimodal Large Language Models (MLLMs) to the task of MKGC introduces significant challenges: (1) the large number of image tokens per entity leads to semantic noise and modality conflicts, and (2) the high computational cost of processing large token inputs. To address these issues, we propose Efficient Lightweight Multimodal Large Language Models (ELMM) for MKGC. ELMM proposes a Multi-view Visual Token Compressor (MVTC) based on multi-head attention mechanism, which adaptively compresses image tokens from both textual and visual views, thereby effectively reducing redundancy while retaining necessary information and avoiding modality conflicts. Additionally, we design an attention pruning strategy to remove redundant attention layers from MLLMs, thereby significantly reducing the inference cost. We further introduce a linear projection to compensate for the performance degradation caused by pruning. Extensive experiments on four benchmark datasets demonstrate that ELMM achieves state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16753
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ELMM: Efficient Lightweight Multimodal Large Language Models for Multimodal Knowledge Graph Completion
Huang, Wei
Li, Peining
Liang, Meiyu
Hou, Xu
Du, Junping
Shao, Yingxia
Ye, Guanhua
Liu, Wu
Lu, Kangkang
Yu, Yang
Artificial Intelligence
68T30
H.3.3
Multimodal Knowledge Graphs (MKGs) extend traditional knowledge graphs by incorporating visual and textual modalities, enabling richer and more expressive entity representations. However, existing MKGs often suffer from incompleteness, which hinder their effectiveness in downstream tasks. Therefore, multimodal knowledge graph completion (MKGC) task is receiving increasing attention. While large language models (LLMs) have shown promise for knowledge graph completion (KGC), their application to the multimodal setting remains underexplored. Moreover, applying Multimodal Large Language Models (MLLMs) to the task of MKGC introduces significant challenges: (1) the large number of image tokens per entity leads to semantic noise and modality conflicts, and (2) the high computational cost of processing large token inputs. To address these issues, we propose Efficient Lightweight Multimodal Large Language Models (ELMM) for MKGC. ELMM proposes a Multi-view Visual Token Compressor (MVTC) based on multi-head attention mechanism, which adaptively compresses image tokens from both textual and visual views, thereby effectively reducing redundancy while retaining necessary information and avoiding modality conflicts. Additionally, we design an attention pruning strategy to remove redundant attention layers from MLLMs, thereby significantly reducing the inference cost. We further introduce a linear projection to compensate for the performance degradation caused by pruning. Extensive experiments on four benchmark datasets demonstrate that ELMM achieves state-of-the-art performance.
title ELMM: Efficient Lightweight Multimodal Large Language Models for Multimodal Knowledge Graph Completion
topic Artificial Intelligence
68T30
H.3.3
url https://arxiv.org/abs/2510.16753