CIRCUS: Circuit Consensus under Uncertainty via Stability Ensembles
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912975750692864 |
|---|---|
| author | Parekh, Swapnil |
| author_facet | Parekh, Swapnil |
| contents | Every mechanistic circuit carries an invisible
asterisk: it reflects not just the model's
computation, but the analyst's choice of
pruning threshold. Change that choice and the
circuit changes, yet current practice treats a
single pruned subgraph as ground truth with
no way to distinguish robust structure from
threshold artifacts. We introduce CIRCUS,
which reframes circuit discovery as a problem
of uncertainty over explanations. CIRCUS
prunes one attribution graph under B
configurations, assigns each edge an empirical
inclusion frequency s(e) in [0,1] measuring
how robustly it survives across the
configuration family, and extracts a consensus
circuit of edges present in every view. This
yields a principled core/contingent/noise
decomposition (analogous to posterior
model-inclusion indicators in Bayesian
variable selection) that separates robust
structure from threshold-sensitive artifacts,
with negligible overhead. On Gemma-2-2B and
Llama-3.2-1B, consensus circuits are 40x
smaller than the union of all configurations
while retaining comparable influence-flow
explanatory power, consistently outperform
influence-ranked and random baselines, and are
confirmed causally relevant by activation
patching. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_00523 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | CIRCUS: Circuit Consensus under Uncertainty via Stability Ensembles Parekh, Swapnil Computation and Language Artificial Intelligence Machine Learning Every mechanistic circuit carries an invisible asterisk: it reflects not just the model's computation, but the analyst's choice of pruning threshold. Change that choice and the circuit changes, yet current practice treats a single pruned subgraph as ground truth with no way to distinguish robust structure from threshold artifacts. We introduce CIRCUS, which reframes circuit discovery as a problem of uncertainty over explanations. CIRCUS prunes one attribution graph under B configurations, assigns each edge an empirical inclusion frequency s(e) in [0,1] measuring how robustly it survives across the configuration family, and extracts a consensus circuit of edges present in every view. This yields a principled core/contingent/noise decomposition (analogous to posterior model-inclusion indicators in Bayesian variable selection) that separates robust structure from threshold-sensitive artifacts, with negligible overhead. On Gemma-2-2B and Llama-3.2-1B, consensus circuits are 40x smaller than the union of all configurations while retaining comparable influence-flow explanatory power, consistently outperform influence-ranked and random baselines, and are confirmed causally relevant by activation patching. |
| title | CIRCUS: Circuit Consensus under Uncertainty via Stability Ensembles |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2603.00523 |