HEAL: Hierarchical Embedding Alignment Loss for Improved Retrieval and Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhattarai, Manish, Barron, Ryan, Eren, Maksim, Vu, Minh, Grantcharov, Vesselin, Boureima, Ismael, Stanev, Valentin, Matuszek, Cynthia, Valtchinov, Vladimir, Rasmussen, Kim, Alexandrov, Boian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916510015946752
author Bhattarai, Manish
Barron, Ryan
Eren, Maksim
Vu, Minh
Grantcharov, Vesselin
Boureima, Ismael
Stanev, Valentin
Matuszek, Cynthia
Valtchinov, Vladimir
Rasmussen, Kim
Alexandrov, Boian
author_facet Bhattarai, Manish
Barron, Ryan
Eren, Maksim
Vu, Minh
Grantcharov, Vesselin
Boureima, Ismael
Stanev, Valentin
Matuszek, Cynthia
Valtchinov, Vladimir
Rasmussen, Kim
Alexandrov, Boian
contents Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external document retrieval to provide domain-specific or up-to-date knowledge. The effectiveness of RAG depends on the relevance of retrieved documents, which is influenced by the semantic alignment of embeddings with the domain's specialized content. Although full fine-tuning can align language models to specific domains, it is computationally intensive and demands substantial data. This paper introduces Hierarchical Embedding Alignment Loss (HEAL), a novel method that leverages hierarchical fuzzy clustering with matrix factorization within contrastive learning to efficiently align LLM embeddings with domain-specific content. HEAL computes level/depth-wise contrastive losses and incorporates hierarchical penalties to align embeddings with the underlying relationships in label hierarchies. This approach enhances retrieval relevance and document classification, effectively reducing hallucinations in LLM outputs. In our experiments, we benchmark and evaluate HEAL across diverse domains, including Healthcare, Material Science, Cyber-security, and Applied Maths.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04661
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HEAL: Hierarchical Embedding Alignment Loss for Improved Retrieval and Representation Learning
Bhattarai, Manish
Barron, Ryan
Eren, Maksim
Vu, Minh
Grantcharov, Vesselin
Boureima, Ismael
Stanev, Valentin
Matuszek, Cynthia
Valtchinov, Vladimir
Rasmussen, Kim
Alexandrov, Boian
Information Retrieval
Artificial Intelligence
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external document retrieval to provide domain-specific or up-to-date knowledge. The effectiveness of RAG depends on the relevance of retrieved documents, which is influenced by the semantic alignment of embeddings with the domain's specialized content. Although full fine-tuning can align language models to specific domains, it is computationally intensive and demands substantial data. This paper introduces Hierarchical Embedding Alignment Loss (HEAL), a novel method that leverages hierarchical fuzzy clustering with matrix factorization within contrastive learning to efficiently align LLM embeddings with domain-specific content. HEAL computes level/depth-wise contrastive losses and incorporates hierarchical penalties to align embeddings with the underlying relationships in label hierarchies. This approach enhances retrieval relevance and document classification, effectively reducing hallucinations in LLM outputs. In our experiments, we benchmark and evaluate HEAL across diverse domains, including Healthcare, Material Science, Cyber-security, and Applied Maths.
title HEAL: Hierarchical Embedding Alignment Loss for Improved Retrieval and Representation Learning
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2412.04661