Saved in:
Bibliographic Details
Main Authors: Graham, Robert, Stevinson, Edward, Richter, Leo, Chia, Alexander, Miller, Joseph, Bloom, Joseph Isaac
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.15735
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910043081801728
author Graham, Robert
Stevinson, Edward
Richter, Leo
Chia, Alexander
Miller, Joseph
Bloom, Joseph Isaac
author_facet Graham, Robert
Stevinson, Edward
Richter, Leo
Chia, Alexander
Miller, Joseph
Bloom, Joseph Isaac
contents Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of generating targeted, linguistically fluent inputs that activate specific latent features or elicit model behaviours. We formalise this approach as context modification and present ContextBench -- a benchmark with tasks assessing core method capabilities and potential safety applications. Our evaluation framework measures both elicitation strength (activation of latent features or behaviours) and linguistic fluency, highlighting how current state-of-the-art methods struggle to balance these objectives. We enhance Evolutionary Prompt Optimisation (EPO) with LLM-assistance and diffusion model inpainting, and demonstrate that these variants achieve state-of-the-art performance in balancing elicitation effectiveness and fluency.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15735
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ContextBench: Modifying Contexts for Targeted Latent Activation
Graham, Robert
Stevinson, Edward
Richter, Leo
Chia, Alexander
Miller, Joseph
Bloom, Joseph Isaac
Artificial Intelligence
Machine Learning
Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of generating targeted, linguistically fluent inputs that activate specific latent features or elicit model behaviours. We formalise this approach as context modification and present ContextBench -- a benchmark with tasks assessing core method capabilities and potential safety applications. Our evaluation framework measures both elicitation strength (activation of latent features or behaviours) and linguistic fluency, highlighting how current state-of-the-art methods struggle to balance these objectives. We enhance Evolutionary Prompt Optimisation (EPO) with LLM-assistance and diffusion model inpainting, and demonstrate that these variants achieve state-of-the-art performance in balancing elicitation effectiveness and fluency.
title ContextBench: Modifying Contexts for Targeted Latent Activation
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.15735