Model-based Subsampling for Knowledge Graph Completion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Xincan, Kamigaito, Hidetaka, Hayashi, Katsuhiko, Watanabe, Taro
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916201940123648
author Feng, Xincan
Kamigaito, Hidetaka
Hayashi, Katsuhiko
Watanabe, Taro
author_facet Feng, Xincan
Kamigaito, Hidetaka
Hayashi, Katsuhiko
Watanabe, Taro
contents Subsampling is effective in Knowledge Graph Embedding (KGE) for reducing overfitting caused by the sparsity in Knowledge Graph (KG) datasets. However, current subsampling approaches consider only frequencies of queries that consist of entities and their relations. Thus, the existing subsampling potentially underestimates the appearance probabilities of infrequent queries even if the frequencies of their entities or relations are high. To address this problem, we propose Model-based Subsampling (MBS) and Mixed Subsampling (MIX) to estimate their appearance probabilities through predictions of KGE models. Evaluation results on datasets FB15k-237, WN18RR, and YAGO3-10 showed that our proposed subsampling methods actually improved the KG completion performances for popular KGE models, RotatE, TransE, HAKE, ComplEx, and DistMult.
format Preprint
id arxiv_https___arxiv_org_abs_2309_09296
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Model-based Subsampling for Knowledge Graph Completion
Feng, Xincan
Kamigaito, Hidetaka
Hayashi, Katsuhiko
Watanabe, Taro
Computation and Language
Artificial Intelligence
Machine Learning
Subsampling is effective in Knowledge Graph Embedding (KGE) for reducing overfitting caused by the sparsity in Knowledge Graph (KG) datasets. However, current subsampling approaches consider only frequencies of queries that consist of entities and their relations. Thus, the existing subsampling potentially underestimates the appearance probabilities of infrequent queries even if the frequencies of their entities or relations are high. To address this problem, we propose Model-based Subsampling (MBS) and Mixed Subsampling (MIX) to estimate their appearance probabilities through predictions of KGE models. Evaluation results on datasets FB15k-237, WN18RR, and YAGO3-10 showed that our proposed subsampling methods actually improved the KG completion performances for popular KGE models, RotatE, TransE, HAKE, ComplEx, and DistMult.
title Model-based Subsampling for Knowledge Graph Completion
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2309.09296