WhyLab: A Causal Audit Framework for Stable Agent Self-Improvement

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Anonymous
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901633898643456
author Anonymous
author_facet Anonymous
contents <p>Self-improving AI agents lack runtime safeguards that prevent evaluation drift, fragile outcome acceptance, and unbounded parameter updates from compounding into catastrophic policy degradation. <strong>WhyLab</strong> introduces a causal audit framework comprising three complementary defenses:</p><ul><li><strong>C1</strong>: Information-theoretic drift detection across evaluation streams</li><li><strong>C2</strong>: E-value × Robustness Value dual-threshold filter for fragile outcomes</li><li><strong>C3</strong>: Lyapunov-bounded adaptive damping with observable energy proxy</li></ul><p>Experiments on synthetic environments demonstrate that C1 improves within-horizon detection reliability, C2 substantially reduces fragile acceptance rates, and C3 achieves the lowest violation frequency with strong proxy–state alignment.</p><p>Code: https://github.com/neogenesislab/WhyLab-NeurIPS2026</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18948929
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle WhyLab: A Causal Audit Framework for Stable Agent Self-Improvement
Anonymous
Causal Inference
Agent Self-Improvement
Drift Detection
Sensitivity Analysis
Lyapunov Stability
E-value
Robustness Value
NeurIPS 2026
<p>Self-improving AI agents lack runtime safeguards that prevent evaluation drift, fragile outcome acceptance, and unbounded parameter updates from compounding into catastrophic policy degradation. <strong>WhyLab</strong> introduces a causal audit framework comprising three complementary defenses:</p><ul><li><strong>C1</strong>: Information-theoretic drift detection across evaluation streams</li><li><strong>C2</strong>: E-value × Robustness Value dual-threshold filter for fragile outcomes</li><li><strong>C3</strong>: Lyapunov-bounded adaptive damping with observable energy proxy</li></ul><p>Experiments on synthetic environments demonstrate that C1 improves within-horizon detection reliability, C2 substantially reduces fragile acceptance rates, and C3 achieves the lowest violation frequency with strong proxy–state alignment.</p><p>Code: https://github.com/neogenesislab/WhyLab-NeurIPS2026</p>
title WhyLab: A Causal Audit Framework for Stable Agent Self-Improvement
topic Causal Inference
Agent Self-Improvement
Drift Detection
Sensitivity Analysis
Lyapunov Stability
E-value
Robustness Value
NeurIPS 2026
url https://doi.org/10.5281/zenodo.18948929