Step-Level Sparse Autoencoder for Reasoning Process Interpretation
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Xuan, Liu, Jiayu, Lai, Yuhang, Xu, Hao, Huang, Zhenya, Miao, Ning |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Deep Thinking by Markov Chain of Continuous Thoughts
di: Liu, Jiayu, et al.
Pubblicazione: (2025)
di: Liu, Jiayu, et al.
Pubblicazione: (2025)
Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
di: Zhao, Daniel, et al.
Pubblicazione: (2025)
di: Zhao, Daniel, et al.
Pubblicazione: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
di: Lai, Yuhang, et al.
Pubblicazione: (2026)
di: Lai, Yuhang, et al.
Pubblicazione: (2026)
Interpretable Reward Model via Sparse Autoencoder
di: Zhang, Shuyi, et al.
Pubblicazione: (2025)
di: Zhang, Shuyi, et al.
Pubblicazione: (2025)
Interpreting Attention Layer Outputs with Sparse Autoencoders
di: Kissane, Connor, et al.
Pubblicazione: (2024)
di: Kissane, Connor, et al.
Pubblicazione: (2024)
Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
di: Parsan, Nithin, et al.
Pubblicazione: (2025)
di: Parsan, Nithin, et al.
Pubblicazione: (2025)
Interpreting CFD Surrogates through Sparse Autoencoders
di: Hu, Yeping, et al.
Pubblicazione: (2025)
di: Hu, Yeping, et al.
Pubblicazione: (2025)
Route Sparse Autoencoder to Interpret Large Language Models
di: Shi, Wei, et al.
Pubblicazione: (2025)
di: Shi, Wei, et al.
Pubblicazione: (2025)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
di: Makelov, Aleksandar, et al.
Pubblicazione: (2024)
di: Makelov, Aleksandar, et al.
Pubblicazione: (2024)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025)
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025)
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
di: Paek, Nathan, et al.
Pubblicazione: (2025)
di: Paek, Nathan, et al.
Pubblicazione: (2025)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
di: Erdogan, Ege, et al.
Pubblicazione: (2025)
di: Erdogan, Ege, et al.
Pubblicazione: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
di: Marks, Luke, et al.
Pubblicazione: (2024)
di: Marks, Luke, et al.
Pubblicazione: (2024)
Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
di: Ye, Mengyu, et al.
Pubblicazione: (2025)
di: Ye, Mengyu, et al.
Pubblicazione: (2025)
Interpretable Company Similarity with Sparse Autoencoders
di: Molinari, Marco, et al.
Pubblicazione: (2024)
di: Molinari, Marco, et al.
Pubblicazione: (2024)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
di: Kurochkin, Vadim, et al.
Pubblicazione: (2025)
di: Kurochkin, Vadim, et al.
Pubblicazione: (2025)
Interpreting and Steering Protein Language Models through Sparse Autoencoders
di: Garcia, Edith Natalia Villegas, et al.
Pubblicazione: (2025)
di: Garcia, Edith Natalia Villegas, et al.
Pubblicazione: (2025)
Interpreting CLIP with Hierarchical Sparse Autoencoders
di: Zaigrajew, Vladimir, et al.
Pubblicazione: (2025)
di: Zaigrajew, Vladimir, et al.
Pubblicazione: (2025)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
di: Tolooshams, Bahareh, et al.
Pubblicazione: (2025)
di: Tolooshams, Bahareh, et al.
Pubblicazione: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
di: O'Neill, Charles, et al.
Pubblicazione: (2025)
di: O'Neill, Charles, et al.
Pubblicazione: (2025)
XNNTab -- Interpretable Neural Networks for Tabular Data using Sparse Autoencoders
di: Elhadri, Khawla, et al.
Pubblicazione: (2025)
di: Elhadri, Khawla, et al.
Pubblicazione: (2025)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
di: Ma, George, et al.
Pubblicazione: (2026)
di: Ma, George, et al.
Pubblicazione: (2026)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
di: Bussmann, Bart, et al.
Pubblicazione: (2025)
di: Bussmann, Bart, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
di: Tahimic, Kriz, et al.
Pubblicazione: (2025)
di: Tahimic, Kriz, et al.
Pubblicazione: (2025)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
di: Jiang, Nick, et al.
Pubblicazione: (2025)
di: Jiang, Nick, et al.
Pubblicazione: (2025)
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
di: Yao, Yifei, et al.
Pubblicazione: (2025)
di: Yao, Yifei, et al.
Pubblicazione: (2025)
Sparse Autoencoders for Interpretable Medical Image Representation Learning
di: Wesp, Philipp, et al.
Pubblicazione: (2026)
di: Wesp, Philipp, et al.
Pubblicazione: (2026)
SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation
di: Lu, Zhenyu, et al.
Pubblicazione: (2026)
di: Lu, Zhenyu, et al.
Pubblicazione: (2026)
Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrieval
di: Park, Seongwan, et al.
Pubblicazione: (2025)
di: Park, Seongwan, et al.
Pubblicazione: (2025)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
di: Cywiński, Bartosz, et al.
Pubblicazione: (2025)
di: Cywiński, Bartosz, et al.
Pubblicazione: (2025)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
Ensembling Sparse Autoencoders
di: Gadgil, Soham, et al.
Pubblicazione: (2025)
di: Gadgil, Soham, et al.
Pubblicazione: (2025)
Stabilizing Efficient Reasoning with Step-Level Advantage Selection
di: Wang, Han, et al.
Pubblicazione: (2026)
di: Wang, Han, et al.
Pubblicazione: (2026)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
di: Klenitskiy, Anton, et al.
Pubblicazione: (2025)
di: Klenitskiy, Anton, et al.
Pubblicazione: (2025)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
di: Thasarathan, Harrish, et al.
Pubblicazione: (2025)
di: Thasarathan, Harrish, et al.
Pubblicazione: (2025)
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
di: Roy, Dip, et al.
Pubblicazione: (2025)
di: Roy, Dip, et al.
Pubblicazione: (2025)
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
di: Yeung, Calvin, et al.
Pubblicazione: (2026)
di: Yeung, Calvin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Deep Thinking by Markov Chain of Continuous Thoughts
di: Liu, Jiayu, et al.
Pubblicazione: (2025) -
Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
di: Zhao, Daniel, et al.
Pubblicazione: (2025) -
Transcoders Beat Sparse Autoencoders for Interpretability
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025) -
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
di: Lai, Yuhang, et al.
Pubblicazione: (2026) -
Interpretable Reward Model via Sparse Autoencoder
di: Zhang, Shuyi, et al.
Pubblicazione: (2025)