WhyLab: A Causal Audit Framework for Stable Agent Self-Improvement
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901633898643456 |
|---|---|
| author | Anonymous |
| author_facet | Anonymous |
| contents | <p>Self-improving AI agents lack runtime safeguards that prevent evaluation drift, fragile outcome acceptance, and unbounded parameter updates from compounding into catastrophic policy degradation. <strong>WhyLab</strong> introduces a causal audit framework comprising three complementary defenses:</p><ul><li><strong>C1</strong>: Information-theoretic drift detection across evaluation streams</li><li><strong>C2</strong>: E-value × Robustness Value dual-threshold filter for fragile outcomes</li><li><strong>C3</strong>: Lyapunov-bounded adaptive damping with observable energy proxy</li></ul><p>Experiments on synthetic environments demonstrate that C1 improves within-horizon detection reliability, C2 substantially reduces fragile acceptance rates, and C3 achieves the lowest violation frequency with strong proxy–state alignment.</p><p>Code: https://github.com/neogenesislab/WhyLab-NeurIPS2026</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18948929 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | WhyLab: A Causal Audit Framework for Stable Agent Self-Improvement Anonymous Causal Inference Agent Self-Improvement Drift Detection Sensitivity Analysis Lyapunov Stability E-value Robustness Value NeurIPS 2026 <p>Self-improving AI agents lack runtime safeguards that prevent evaluation drift, fragile outcome acceptance, and unbounded parameter updates from compounding into catastrophic policy degradation. <strong>WhyLab</strong> introduces a causal audit framework comprising three complementary defenses:</p><ul><li><strong>C1</strong>: Information-theoretic drift detection across evaluation streams</li><li><strong>C2</strong>: E-value × Robustness Value dual-threshold filter for fragile outcomes</li><li><strong>C3</strong>: Lyapunov-bounded adaptive damping with observable energy proxy</li></ul><p>Experiments on synthetic environments demonstrate that C1 improves within-horizon detection reliability, C2 substantially reduces fragile acceptance rates, and C3 achieves the lowest violation frequency with strong proxy–state alignment.</p><p>Code: https://github.com/neogenesislab/WhyLab-NeurIPS2026</p> |
| title | WhyLab: A Causal Audit Framework for Stable Agent Self-Improvement |
| topic | Causal Inference Agent Self-Improvement Drift Detection Sensitivity Analysis Lyapunov Stability E-value Robustness Value NeurIPS 2026 |
| url | https://doi.org/10.5281/zenodo.18948929 |