Saved in:
Bibliographic Details
Main Authors: Cardoso, João N., Oliveira, Arlindo L., Martins, Bruno
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.17867
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910027866963968
author Cardoso, João N.
Oliveira, Arlindo L.
Martins, Bruno
author_facet Cardoso, João N.
Oliveira, Arlindo L.
Martins, Bruno
contents Understanding what features are encoded by learned directions in LLM activation space requires identifying inputs that strongly activate them. Feature visualization, which optimizes inputs to maximally activate a target direction, offers an alternative to costly dataset search approaches, but remains underexplored for LLMs due to the discrete nature of text. Furthermore, existing prompt optimization techniques are poorly suited to this domain, which is highly prone to local minima. To overcome these limitations, we introduce ADAPT, a hybrid method combining beam search initialization with adaptive gradient-guided mutation, designed around these failure modes. We evaluate on Sparse Autoencoder latents from Gemma 2 2B, proposing metrics grounded in dataset activation statistics to enable rigorous comparison, and show that ADAPT consistently outperforms prior methods across layers and latent types. Our results establish that feature visualization for LLMs is tractable, but requires design assumptions tailored to the domain.
format Preprint
id arxiv_https___arxiv_org_abs_2602_17867
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ADAPT: Hybrid Prompt Optimization for LLM Feature Visualization
Cardoso, João N.
Oliveira, Arlindo L.
Martins, Bruno
Machine Learning
Computation and Language
Understanding what features are encoded by learned directions in LLM activation space requires identifying inputs that strongly activate them. Feature visualization, which optimizes inputs to maximally activate a target direction, offers an alternative to costly dataset search approaches, but remains underexplored for LLMs due to the discrete nature of text. Furthermore, existing prompt optimization techniques are poorly suited to this domain, which is highly prone to local minima. To overcome these limitations, we introduce ADAPT, a hybrid method combining beam search initialization with adaptive gradient-guided mutation, designed around these failure modes. We evaluate on Sparse Autoencoder latents from Gemma 2 2B, proposing metrics grounded in dataset activation statistics to enable rigorous comparison, and show that ADAPT consistently outperforms prior methods across layers and latent types. Our results establish that feature visualization for LLMs is tractable, but requires design assumptions tailored to the domain.
title ADAPT: Hybrid Prompt Optimization for LLM Feature Visualization
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2602.17867