Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Shun, Loakman, Tyler, Lei, Youbo, Liu, Yi, Yang, Bohao, Zhao, Yuting, Yang, Dong, Lin, Chenghua
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911046811254784
author Wang, Shun
Loakman, Tyler
Lei, Youbo
Liu, Yi
Yang, Bohao
Zhao, Yuting
Yang, Dong
Lin, Chenghua
author_facet Wang, Shun
Loakman, Tyler
Lei, Youbo
Liu, Yi
Yang, Bohao
Zhao, Yuting
Yang, Dong
Lin, Chenghua
contents Large Language Models (LLMs) are traditionally viewed as black-box algorithms, therefore reducing trustworthiness and obscuring potential approaches to increasing performance on downstream tasks. In this work, we apply an effective LLM decomposition method using a dictionary-learning approach with sparse autoencoders. This helps extract monosemantic features from polysemantic LLM neurons. Remarkably, our work identifies model-internal misunderstanding, allowing the automatic reformulation of the prompts with additional annotations to improve the interpretation by LLMs. Moreover, this approach demonstrates a significant performance improvement in downstream tasks, such as mathematical reasoning and metaphor detection.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06427
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
Wang, Shun
Loakman, Tyler
Lei, Youbo
Liu, Yi
Yang, Bohao
Zhao, Yuting
Yang, Dong
Lin, Chenghua
Computation and Language
Machine Learning
Large Language Models (LLMs) are traditionally viewed as black-box algorithms, therefore reducing trustworthiness and obscuring potential approaches to increasing performance on downstream tasks. In this work, we apply an effective LLM decomposition method using a dictionary-learning approach with sparse autoencoders. This helps extract monosemantic features from polysemantic LLM neurons. Remarkably, our work identifies model-internal misunderstanding, allowing the automatic reformulation of the prompts with additional annotations to improve the interpretation by LLMs. Moreover, this approach demonstrates a significant performance improvement in downstream tasks, such as mathematical reasoning and metaphor detection.
title Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.06427