Domain Specific Data Distillation and Multi-modal Embedding Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peddiraju, Sharadind, Rajagopal, Srini
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912088155226112
author Peddiraju, Sharadind
Rajagopal, Srini
author_facet Peddiraju, Sharadind
Rajagopal, Srini
contents The challenge of creating domain-centric embeddings arises from the abundance of unstructured data and the scarcity of domain-specific structured data. Conventional embedding techniques often rely on either modality, limiting their applicability and efficacy. This paper introduces a novel modeling approach that leverages structured data to filter noise from unstructured data, resulting in embeddings with high precision and recall for domain-specific attribute prediction. The proposed model operates within a Hybrid Collaborative Filtering (HCF) framework, where generic entity representations are fine-tuned through relevant item prediction tasks. Our experiments, focusing on the cloud computing domain, demonstrate that HCF-based embeddings outperform AutoEncoder-based embeddings (using purely unstructured data), achieving a 28% lift in precision and an 11% lift in recall for domain-specific attribute prediction.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20325
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Domain Specific Data Distillation and Multi-modal Embedding Generation
Peddiraju, Sharadind
Rajagopal, Srini
Machine Learning
Social and Information Networks
I.2.4; H.3.3; I.5.3; I.1.2; H.2.5
The challenge of creating domain-centric embeddings arises from the abundance of unstructured data and the scarcity of domain-specific structured data. Conventional embedding techniques often rely on either modality, limiting their applicability and efficacy. This paper introduces a novel modeling approach that leverages structured data to filter noise from unstructured data, resulting in embeddings with high precision and recall for domain-specific attribute prediction. The proposed model operates within a Hybrid Collaborative Filtering (HCF) framework, where generic entity representations are fine-tuned through relevant item prediction tasks. Our experiments, focusing on the cloud computing domain, demonstrate that HCF-based embeddings outperform AutoEncoder-based embeddings (using purely unstructured data), achieving a 28% lift in precision and an 11% lift in recall for domain-specific attribute prediction.
title Domain Specific Data Distillation and Multi-modal Embedding Generation
topic Machine Learning
Social and Information Networks
I.2.4; H.3.3; I.5.3; I.1.2; H.2.5
url https://arxiv.org/abs/2410.20325