Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.12421 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911209763110912 |
|---|---|
| author | La Malfa, Emanuele Vadillo, Jon Molinari, Marco Wooldridge, Michael |
| author_facet | La Malfa, Emanuele Vadillo, Jon Molinari, Marco Wooldridge, Michael |
| contents | This paper introduces a formal notion of fixed point explanations, inspired by the "why regress" principle, to assess, through recursive applications, the stability of the interplay between a model and its explainer. Fixed point explanations satisfy properties like minimality, stability, and faithfulness, revealing hidden model behaviours and explanatory weaknesses. We define convergence conditions for several classes of explainers, from feature-based to mechanistic tools like Sparse AutoEncoders, and we report quantitative and qualitative results for several datasets and models, including LLMs such as Llama-3.3-70B. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_12421 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Fixed Point Explainability La Malfa, Emanuele Vadillo, Jon Molinari, Marco Wooldridge, Michael Machine Learning Artificial Intelligence This paper introduces a formal notion of fixed point explanations, inspired by the "why regress" principle, to assess, through recursive applications, the stability of the interplay between a model and its explainer. Fixed point explanations satisfy properties like minimality, stability, and faithfulness, revealing hidden model behaviours and explanatory weaknesses. We define convergence conditions for several classes of explainers, from feature-based to mechanistic tools like Sparse AutoEncoders, and we report quantitative and qualitative results for several datasets and models, including LLMs such as Llama-3.3-70B. |
| title | Fixed Point Explainability |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2505.12421 |