Unveiling and Manipulating Prompt Influence in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Zijian, Zhou, Hanzhang, Zhu, Zixiao, Qian, Junlang, Mao, Kezhi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913356994052096
author Feng, Zijian
Zhou, Hanzhang
Zhu, Zixiao
Qian, Junlang
Mao, Kezhi
author_facet Feng, Zijian
Zhou, Hanzhang
Zhu, Zixiao
Qian, Junlang
Mao, Kezhi
contents Prompts play a crucial role in guiding the responses of Large Language Models (LLMs). However, the intricate role of individual tokens in prompts, known as input saliency, in shaping the responses remains largely underexplored. Existing saliency methods either misalign with LLM generation objectives or rely heavily on linearity assumptions, leading to potential inaccuracies. To address this, we propose Token Distribution Dynamics (TDD), a \textcolor{black}{simple yet effective} approach to unveil and manipulate the role of prompts in generating LLM outputs. TDD leverages the robust interpreting capabilities of the language model head (LM head) to assess input saliency. It projects input tokens into the embedding space and then estimates their significance based on distribution dynamics over the vocabulary. We introduce three TDD variants: forward, backward, and bidirectional, each offering unique insights into token relevance. Extensive experiments reveal that the TDD surpasses state-of-the-art baselines with a big margin in elucidating the causal relationships between prompts and LLM outputs. Beyond mere interpretation, we apply TDD to two prompt manipulation tasks for controlled text generation: zero-shot toxic language suppression and sentiment steering. Empirical results underscore TDD's proficiency in identifying both toxic and sentimental cues in prompts, subsequently mitigating toxicity or modulating sentiment in the generated content.
format Preprint
id arxiv_https___arxiv_org_abs_2405_11891
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unveiling and Manipulating Prompt Influence in Large Language Models
Feng, Zijian
Zhou, Hanzhang
Zhu, Zixiao
Qian, Junlang
Mao, Kezhi
Computation and Language
Artificial Intelligence
Prompts play a crucial role in guiding the responses of Large Language Models (LLMs). However, the intricate role of individual tokens in prompts, known as input saliency, in shaping the responses remains largely underexplored. Existing saliency methods either misalign with LLM generation objectives or rely heavily on linearity assumptions, leading to potential inaccuracies. To address this, we propose Token Distribution Dynamics (TDD), a \textcolor{black}{simple yet effective} approach to unveil and manipulate the role of prompts in generating LLM outputs. TDD leverages the robust interpreting capabilities of the language model head (LM head) to assess input saliency. It projects input tokens into the embedding space and then estimates their significance based on distribution dynamics over the vocabulary. We introduce three TDD variants: forward, backward, and bidirectional, each offering unique insights into token relevance. Extensive experiments reveal that the TDD surpasses state-of-the-art baselines with a big margin in elucidating the causal relationships between prompts and LLM outputs. Beyond mere interpretation, we apply TDD to two prompt manipulation tasks for controlled text generation: zero-shot toxic language suppression and sentiment steering. Empirical results underscore TDD's proficiency in identifying both toxic and sentimental cues in prompts, subsequently mitigating toxicity or modulating sentiment in the generated content.
title Unveiling and Manipulating Prompt Influence in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.11891