Domain Specific Data Distillation and Multi-modal Embedding Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912088155226112 |
|---|---|
| author | Peddiraju, Sharadind Rajagopal, Srini |
| author_facet | Peddiraju, Sharadind Rajagopal, Srini |
| contents | The challenge of creating domain-centric embeddings arises from the abundance of unstructured data and the scarcity of domain-specific structured data. Conventional embedding techniques often rely on either modality, limiting their applicability and efficacy. This paper introduces a novel modeling approach that leverages structured data to filter noise from unstructured data, resulting in embeddings with high precision and recall for domain-specific attribute prediction. The proposed model operates within a Hybrid Collaborative Filtering (HCF) framework, where generic entity representations are fine-tuned through relevant item prediction tasks. Our experiments, focusing on the cloud computing domain, demonstrate that HCF-based embeddings outperform AutoEncoder-based embeddings (using purely unstructured data), achieving a 28% lift in precision and an 11% lift in recall for domain-specific attribute prediction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_20325 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Domain Specific Data Distillation and Multi-modal Embedding Generation Peddiraju, Sharadind Rajagopal, Srini Machine Learning Social and Information Networks I.2.4; H.3.3; I.5.3; I.1.2; H.2.5 The challenge of creating domain-centric embeddings arises from the abundance of unstructured data and the scarcity of domain-specific structured data. Conventional embedding techniques often rely on either modality, limiting their applicability and efficacy. This paper introduces a novel modeling approach that leverages structured data to filter noise from unstructured data, resulting in embeddings with high precision and recall for domain-specific attribute prediction. The proposed model operates within a Hybrid Collaborative Filtering (HCF) framework, where generic entity representations are fine-tuned through relevant item prediction tasks. Our experiments, focusing on the cloud computing domain, demonstrate that HCF-based embeddings outperform AutoEncoder-based embeddings (using purely unstructured data), achieving a 28% lift in precision and an 11% lift in recall for domain-specific attribute prediction. |
| title | Domain Specific Data Distillation and Multi-modal Embedding Generation |
| topic | Machine Learning Social and Information Networks I.2.4; H.3.3; I.5.3; I.1.2; H.2.5 |
| url | https://arxiv.org/abs/2410.20325 |