Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Deng, Boyi, Wang, Xu, Wang, Yaoning, Wan, Yu, Ma, Yubo, Yang, Baosong, Wei, Haoran, Tang, Jialong, Lin, Huan, Gao, Ruize, Li, Tianhao, Cao, Qian, Ren, Xuancheng, Deng, Xiaodong, Yang, An, Huang, Fei, Liu, Dayiheng, Zhou, Jingren
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914557940727808
author Deng, Boyi
Wang, Xu
Wang, Yaoning
Wan, Yu
Ma, Yubo
Yang, Baosong
Wei, Haoran
Tang, Jialong
Lin, Huan
Gao, Ruize
Li, Tianhao
Cao, Qian
Ren, Xuancheng
Deng, Xiaodong
Yang, An
Huang, Fei
Liu, Dayiheng
Zhou, Jingren
author_facet Deng, Boyi
Wang, Xu
Wang, Yaoning
Wan, Yu
Ma, Yubo
Yang, Baosong
Wei, Haoran
Tang, Jialong
Lin, Huan
Gao, Ruize
Li, Tianhao
Cao, Qian
Ren, Xuancheng
Deng, Xiaodong
Yang, An
Huang, Fei
Liu, Dayiheng
Zhou, Jingren
contents Large language models have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque, limiting our ability to inspect, control, and systematically improve them. This opacity motivates a growing body of research in mechanistic interpretability, with sparse autoencoders (SAEs) emerging as one of the most promising tools for decomposing model activations into sparse, interpretable feature representations. We introduce Qwen-Scope, an open-source suite of SAEs built on the Qwen model family, comprising 14 groups of SAEs across 7 model variants from the Qwen3 and Qwen3.5 series, covering both dense and mixture-of-expert architectures. Built on top of these SAEs, we show that SAEs can go beyond post-hoc analysis to serve as practical interfaces for model development along four directions: (i) inference-time steering, where SAE feature directions control language, concepts, and preferences without modifying model weights; (ii) evaluation analysis, where activated SAE features provide a representation-level proxy for benchmark redundancy and capability coverage; (iii) data-centric workflows, where SAE features support multilingual toxicity classification and safety-oriented data synthesis; and (iv) post-training optimization, where SAE-derived signals are incorporated into supervised fine-tuning and reinforcement learning objectives to mitigate undesirable behaviors such as code-switching and repetition. Together, these results demonstrate that SAEs can serve not only as post-hoc analysis tools, but also as reusable representation-level interfaces for diagnosing, controlling, evaluating, and improving large language models. By open-sourcing Qwen-Scope, we aim to support mechanistic research and accelerate practical workflows that connect model internals to downstream behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11887
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
Deng, Boyi
Wang, Xu
Wang, Yaoning
Wan, Yu
Ma, Yubo
Yang, Baosong
Wei, Haoran
Tang, Jialong
Lin, Huan
Gao, Ruize
Li, Tianhao
Cao, Qian
Ren, Xuancheng
Deng, Xiaodong
Yang, An
Huang, Fei
Liu, Dayiheng
Zhou, Jingren
Computation and Language
Machine Learning
Large language models have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque, limiting our ability to inspect, control, and systematically improve them. This opacity motivates a growing body of research in mechanistic interpretability, with sparse autoencoders (SAEs) emerging as one of the most promising tools for decomposing model activations into sparse, interpretable feature representations. We introduce Qwen-Scope, an open-source suite of SAEs built on the Qwen model family, comprising 14 groups of SAEs across 7 model variants from the Qwen3 and Qwen3.5 series, covering both dense and mixture-of-expert architectures. Built on top of these SAEs, we show that SAEs can go beyond post-hoc analysis to serve as practical interfaces for model development along four directions: (i) inference-time steering, where SAE feature directions control language, concepts, and preferences without modifying model weights; (ii) evaluation analysis, where activated SAE features provide a representation-level proxy for benchmark redundancy and capability coverage; (iii) data-centric workflows, where SAE features support multilingual toxicity classification and safety-oriented data synthesis; and (iv) post-training optimization, where SAE-derived signals are incorporated into supervised fine-tuning and reinforcement learning objectives to mitigate undesirable behaviors such as code-switching and repetition. Together, these results demonstrate that SAEs can serve not only as post-hoc analysis tools, but also as reusable representation-level interfaces for diagnosing, controlling, evaluating, and improving large language models. By open-sourcing Qwen-Scope, we aim to support mechanistic research and accelerate practical workflows that connect model internals to downstream behavior.
title Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2605.11887