Enhancing Contextual Understanding in Large Language Models through Contrastive Decoding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhao, Zheng, Monti, Emilio, Lehmann, Jens, Assem, Haytham
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911866864795648
author Zhao, Zheng
Monti, Emilio
Lehmann, Jens
Assem, Haytham
author_facet Zhao, Zheng
Monti, Emilio
Lehmann, Jens
Assem, Haytham
contents Large language models (LLMs) tend to inadequately integrate input context during text generation, relying excessively on encoded prior knowledge in model parameters, potentially resulting in generated text with factual inconsistencies or contextually unfaithful content. LLMs utilize two primary knowledge sources: 1) prior (parametric) knowledge from pretraining, and 2) contextual (non-parametric) knowledge from input prompts. The study addresses the open question of how LLMs effectively balance these knowledge sources during the generation process, specifically in the context of open-domain question answering. To address this issue, we introduce a novel approach integrating contrastive decoding with adversarial irrelevant passages as negative samples to enhance robust context grounding during generation. Notably, our method operates at inference time without requiring further training. We conduct comprehensive experiments to demonstrate its applicability and effectiveness, providing empirical evidence showcasing its superiority over existing methodologies. Our code is publicly available at: https://github.com/amazon-science/ContextualUnderstanding-ContrastiveDecoding.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02750
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Contextual Understanding in Large Language Models through Contrastive Decoding
Zhao, Zheng
Monti, Emilio
Lehmann, Jens
Assem, Haytham
Computation and Language
Artificial Intelligence
Large language models (LLMs) tend to inadequately integrate input context during text generation, relying excessively on encoded prior knowledge in model parameters, potentially resulting in generated text with factual inconsistencies or contextually unfaithful content. LLMs utilize two primary knowledge sources: 1) prior (parametric) knowledge from pretraining, and 2) contextual (non-parametric) knowledge from input prompts. The study addresses the open question of how LLMs effectively balance these knowledge sources during the generation process, specifically in the context of open-domain question answering. To address this issue, we introduce a novel approach integrating contrastive decoding with adversarial irrelevant passages as negative samples to enhance robust context grounding during generation. Notably, our method operates at inference time without requiring further training. We conduct comprehensive experiments to demonstrate its applicability and effectiveness, providing empirical evidence showcasing its superiority over existing methodologies. Our code is publicly available at: https://github.com/amazon-science/ContextualUnderstanding-ContrastiveDecoding.
title Enhancing Contextual Understanding in Large Language Models through Contrastive Decoding
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.02750