HEAL: Hierarchical Embedding Alignment Loss for Improved Retrieval and Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916510015946752 |
|---|---|
| author | Bhattarai, Manish Barron, Ryan Eren, Maksim Vu, Minh Grantcharov, Vesselin Boureima, Ismael Stanev, Valentin Matuszek, Cynthia Valtchinov, Vladimir Rasmussen, Kim Alexandrov, Boian |
| author_facet | Bhattarai, Manish Barron, Ryan Eren, Maksim Vu, Minh Grantcharov, Vesselin Boureima, Ismael Stanev, Valentin Matuszek, Cynthia Valtchinov, Vladimir Rasmussen, Kim Alexandrov, Boian |
| contents | Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external document retrieval to provide domain-specific or up-to-date knowledge. The effectiveness of RAG depends on the relevance of retrieved documents, which is influenced by the semantic alignment of embeddings with the domain's specialized content. Although full fine-tuning can align language models to specific domains, it is computationally intensive and demands substantial data. This paper introduces Hierarchical Embedding Alignment Loss (HEAL), a novel method that leverages hierarchical fuzzy clustering with matrix factorization within contrastive learning to efficiently align LLM embeddings with domain-specific content. HEAL computes level/depth-wise contrastive losses and incorporates hierarchical penalties to align embeddings with the underlying relationships in label hierarchies. This approach enhances retrieval relevance and document classification, effectively reducing hallucinations in LLM outputs. In our experiments, we benchmark and evaluate HEAL across diverse domains, including Healthcare, Material Science, Cyber-security, and Applied Maths. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_04661 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | HEAL: Hierarchical Embedding Alignment Loss for Improved Retrieval and Representation Learning Bhattarai, Manish Barron, Ryan Eren, Maksim Vu, Minh Grantcharov, Vesselin Boureima, Ismael Stanev, Valentin Matuszek, Cynthia Valtchinov, Vladimir Rasmussen, Kim Alexandrov, Boian Information Retrieval Artificial Intelligence Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external document retrieval to provide domain-specific or up-to-date knowledge. The effectiveness of RAG depends on the relevance of retrieved documents, which is influenced by the semantic alignment of embeddings with the domain's specialized content. Although full fine-tuning can align language models to specific domains, it is computationally intensive and demands substantial data. This paper introduces Hierarchical Embedding Alignment Loss (HEAL), a novel method that leverages hierarchical fuzzy clustering with matrix factorization within contrastive learning to efficiently align LLM embeddings with domain-specific content. HEAL computes level/depth-wise contrastive losses and incorporates hierarchical penalties to align embeddings with the underlying relationships in label hierarchies. This approach enhances retrieval relevance and document classification, effectively reducing hallucinations in LLM outputs. In our experiments, we benchmark and evaluate HEAL across diverse domains, including Healthcare, Material Science, Cyber-security, and Applied Maths. |
| title | HEAL: Hierarchical Embedding Alignment Loss for Improved Retrieval and Representation Learning |
| topic | Information Retrieval Artificial Intelligence |
| url | https://arxiv.org/abs/2412.04661 |