Saved in:
Bibliographic Details
Main Authors: Yu, Haonan, Liu, Junhao, Yan, Zhenyu, Lin, Haoran, Zhang, Xin
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.18474
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914460813230080
author Yu, Haonan
Liu, Junhao
Yan, Zhenyu
Lin, Haoran
Zhang, Xin
author_facet Yu, Haonan
Liu, Junhao
Yan, Zhenyu
Lin, Haoran
Zhang, Xin
contents Precise behavioral control of large language models (LLMs) is critical for complex applications. However, existing methods often incur high training costs, lack natural language controllability, or compromise semantic coherence. To bridge this gap, we propose WASD (unWeaving Actionable Sufficient Directives), a novel framework that explains model behavior by identifying sufficient neural conditions for token generation. Our method represents candidate conditions as neuron-activation predicates and iteratively searches for a minimal set that guarantees the current output under input perturbations. Experiments on SST-2 and CounterFact with the Gemma-2-2B model demonstrate that our approach produces explanations that are more stable, accurate, and concise than conventional attribution graphs. Moreover, through a case study on controlling cross-lingual output generation, we validated the practical effectiveness of WASD in controlling model behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18474
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WASD: Locating Critical Neurons as Sufficient Conditions for Explaining and Controlling LLM Behavior
Yu, Haonan
Liu, Junhao
Yan, Zhenyu
Lin, Haoran
Zhang, Xin
Computation and Language
Artificial Intelligence
Precise behavioral control of large language models (LLMs) is critical for complex applications. However, existing methods often incur high training costs, lack natural language controllability, or compromise semantic coherence. To bridge this gap, we propose WASD (unWeaving Actionable Sufficient Directives), a novel framework that explains model behavior by identifying sufficient neural conditions for token generation. Our method represents candidate conditions as neuron-activation predicates and iteratively searches for a minimal set that guarantees the current output under input perturbations. Experiments on SST-2 and CounterFact with the Gemma-2-2B model demonstrate that our approach produces explanations that are more stable, accurate, and concise than conventional attribution graphs. Moreover, through a case study on controlling cross-lingual output generation, we validated the practical effectiveness of WASD in controlling model behavior.
title WASD: Locating Critical Neurons as Sufficient Conditions for Explaining and Controlling LLM Behavior
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.18474