Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911046811254784 |
|---|---|
| author | Wang, Shun Loakman, Tyler Lei, Youbo Liu, Yi Yang, Bohao Zhao, Yuting Yang, Dong Lin, Chenghua |
| author_facet | Wang, Shun Loakman, Tyler Lei, Youbo Liu, Yi Yang, Bohao Zhao, Yuting Yang, Dong Lin, Chenghua |
| contents | Large Language Models (LLMs) are traditionally viewed as black-box algorithms, therefore reducing trustworthiness and obscuring potential approaches to increasing performance on downstream tasks. In this work, we apply an effective LLM decomposition method using a dictionary-learning approach with sparse autoencoders. This helps extract monosemantic features from polysemantic LLM neurons. Remarkably, our work identifies model-internal misunderstanding, allowing the automatic reformulation of the prompts with additional annotations to improve the interpretation by LLMs. Moreover, this approach demonstrates a significant performance improvement in downstream tasks, such as mathematical reasoning and metaphor detection. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_06427 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders Wang, Shun Loakman, Tyler Lei, Youbo Liu, Yi Yang, Bohao Zhao, Yuting Yang, Dong Lin, Chenghua Computation and Language Machine Learning Large Language Models (LLMs) are traditionally viewed as black-box algorithms, therefore reducing trustworthiness and obscuring potential approaches to increasing performance on downstream tasks. In this work, we apply an effective LLM decomposition method using a dictionary-learning approach with sparse autoencoders. This helps extract monosemantic features from polysemantic LLM neurons. Remarkably, our work identifies model-internal misunderstanding, allowing the automatic reformulation of the prompts with additional annotations to improve the interpretation by LLMs. Moreover, this approach demonstrates a significant performance improvement in downstream tasks, such as mathematical reasoning and metaphor detection. |
| title | Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2507.06427 |