Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Haiyan, He, Zirui, Wang, Guanchu, Payani, Ali, Li, Yingcong, Du, Mengnan
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913161592963072
author Zhao, Haiyan
He, Zirui
Wang, Guanchu
Payani, Ali
Li, Yingcong
Du, Mengnan
author_facet Zhao, Haiyan
He, Zirui
Wang, Guanchu
Payani, Ali
Li, Yingcong
Du, Mengnan
contents Activation verbalization explains hidden representations in natural language, but existing methods are mostly limited to self-explanation, where each model explains only its own activations. We introduce Universal Activation Verbalizer (UAV), a framework that uses a shared decoder to explain activations from heterogeneous donor models. UAV learns a lightweight adapter that converts donor activations into soft tokens in decoder's embedding space, and further supports adapter-only transfer by reusing a frozen decoder-side LoRA while training only a new adapter for another donor. Across classification, fact retrieval, and gist summarization, UAV remains competitive with strong self-explanation baselines while enabling cross-model verbalization across model families and scales. Ablations show that decoder-side tuning mainly improves task behavior, whereas the adapter provides the activation-grounded factual and semantic information needed for faithful explanations.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25903
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
Zhao, Haiyan
He, Zirui
Wang, Guanchu
Payani, Ali
Li, Yingcong
Du, Mengnan
Computation and Language
Machine Learning
Activation verbalization explains hidden representations in natural language, but existing methods are mostly limited to self-explanation, where each model explains only its own activations. We introduce Universal Activation Verbalizer (UAV), a framework that uses a shared decoder to explain activations from heterogeneous donor models. UAV learns a lightweight adapter that converts donor activations into soft tokens in decoder's embedding space, and further supports adapter-only transfer by reusing a frozen decoder-side LoRA while training only a new adapter for another donor. Across classification, fact retrieval, and gist summarization, UAV remains competitive with strong self-explanation baselines while enabling cross-model verbalization across model families and scales. Ablations show that decoder-side tuning mainly improves task behavior, whereas the adapter provides the activation-grounded factual and semantic information needed for faithful explanations.
title Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2605.25903