CompanyKG: A Large-Scale Heterogeneous Graph for Company Similarity Quantification
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914829707509760 |
|---|---|
| author | Cao, Lele von Ehrenheim, Vilhelm Granroth-Wilding, Mark Stahl, Richard Anselmo McCornack, Andrew Catovic, Armin Rocha, Dhiana Deva Cavacanti |
| author_facet | Cao, Lele von Ehrenheim, Vilhelm Granroth-Wilding, Mark Stahl, Richard Anselmo McCornack, Andrew Catovic, Armin Rocha, Dhiana Deva Cavacanti |
| contents | In the investment industry, it is often essential to carry out fine-grained company similarity quantification for a range of purposes, including market mapping, competitor analysis, and mergers and acquisitions. We propose and publish a knowledge graph, named CompanyKG, to represent and learn diverse company features and relations. Specifically, 1.17 million companies are represented as nodes enriched with company description embeddings; and 15 different inter-company relations result in 51.06 million weighted edges. To enable a comprehensive assessment of methods for company similarity quantification, we have devised and compiled three evaluation tasks with annotated test sets: similarity prediction, competitor retrieval and similarity ranking. We present extensive benchmarking results for 11 reproducible predictive methods categorized into three groups: node-only, edge-only, and node+edge. To the best of our knowledge, CompanyKG is the first large-scale heterogeneous graph dataset originating from a real-world investment platform, tailored for quantifying inter-company similarity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2306_10649 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | CompanyKG: A Large-Scale Heterogeneous Graph for Company Similarity Quantification Cao, Lele von Ehrenheim, Vilhelm Granroth-Wilding, Mark Stahl, Richard Anselmo McCornack, Andrew Catovic, Armin Rocha, Dhiana Deva Cavacanti Artificial Intelligence Computational Engineering, Finance, and Science Databases Machine Learning 05C85, 05C12, 68T07, 68T50, 05C90 E.0; I.2.1; I.2.6; H.4.0; J.0; I.2.8; I.2.7 In the investment industry, it is often essential to carry out fine-grained company similarity quantification for a range of purposes, including market mapping, competitor analysis, and mergers and acquisitions. We propose and publish a knowledge graph, named CompanyKG, to represent and learn diverse company features and relations. Specifically, 1.17 million companies are represented as nodes enriched with company description embeddings; and 15 different inter-company relations result in 51.06 million weighted edges. To enable a comprehensive assessment of methods for company similarity quantification, we have devised and compiled three evaluation tasks with annotated test sets: similarity prediction, competitor retrieval and similarity ranking. We present extensive benchmarking results for 11 reproducible predictive methods categorized into three groups: node-only, edge-only, and node+edge. To the best of our knowledge, CompanyKG is the first large-scale heterogeneous graph dataset originating from a real-world investment platform, tailored for quantifying inter-company similarity. |
| title | CompanyKG: A Large-Scale Heterogeneous Graph for Company Similarity Quantification |
| topic | Artificial Intelligence Computational Engineering, Finance, and Science Databases Machine Learning 05C85, 05C12, 68T07, 68T50, 05C90 E.0; I.2.1; I.2.6; H.4.0; J.0; I.2.8; I.2.7 |
| url | https://arxiv.org/abs/2306.10649 |