Do Sparse Autoencoders Identify Reasoning Features in Language Models?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, George, Liang, Zhongyuan, Chen, Irene Y., Sojoudi, Somayeh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
von: Pfrommer, Samuel, et al.
Veröffentlicht: (2025)
von: Pfrommer, Samuel, et al.
Veröffentlicht: (2025)
Revealing Treatment Non-Adherence Bias in Clinical Machine Learning Using Large Language Models
von: Liang, Zhongyuan, et al.
Veröffentlicht: (2025)
von: Liang, Zhongyuan, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Flow-Matching Policies
von: Pfrommer, Samuel, et al.
Veröffentlicht: (2025)
von: Pfrommer, Samuel, et al.
Veröffentlicht: (2025)
Absence of spurious solutions far from ground truth: A low-rank analysis with high-order losses
von: Ma, Ziye, et al.
Veröffentlicht: (2024)
von: Ma, Ziye, et al.
Veröffentlicht: (2024)
Transport of Algebraic Structure to Latent Embeddings
von: Pfrommer, Samuel, et al.
Veröffentlicht: (2024)
von: Pfrommer, Samuel, et al.
Veröffentlicht: (2024)
Toward Identifiable Sparse Autoencoders
von: Nelson, Walter, et al.
Veröffentlicht: (2026)
von: Nelson, Walter, et al.
Veröffentlicht: (2026)
Route Sparse Autoencoder to Interpret Large Language Models
von: Shi, Wei, et al.
Veröffentlicht: (2025)
von: Shi, Wei, et al.
Veröffentlicht: (2025)
Towards Optimal Branching of Linear and Semidefinite Relaxations for Neural Network Robustness Certification
von: Anderson, Brendon G., et al.
Veröffentlicht: (2021)
von: Anderson, Brendon G., et al.
Veröffentlicht: (2021)
Pausing Policy Learning in Non-stationary Reinforcement Learning
von: Lee, Hyunin, et al.
Veröffentlicht: (2024)
von: Lee, Hyunin, et al.
Veröffentlicht: (2024)
Reinforcement Learning via Value Gradient Flow
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
DRAGON: Distributional Rewards Optimize Diffusion Generative Models
von: Bai, Yatong, et al.
Veröffentlicht: (2025)
von: Bai, Yatong, et al.
Veröffentlicht: (2025)
Steering Language Model Refusal with Sparse Autoencoders
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders
von: Wu, John F., et al.
Veröffentlicht: (2025)
von: Wu, John F., et al.
Veröffentlicht: (2025)
SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
Mixing Classifiers to Alleviate the Accuracy-Robustness Trade-Off
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
von: Joshi, Shruti, et al.
Veröffentlicht: (2025)
von: Joshi, Shruti, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
Sparse Autoencoder Features for Classifications and Transferability
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
von: Cheng, Ziheng, et al.
Veröffentlicht: (2025)
von: Cheng, Ziheng, et al.
Veröffentlicht: (2025)
Do Sparse Autoencoders Generalize? A Case Study of Answerability
von: Heindrich, Lovis, et al.
Veröffentlicht: (2025)
von: Heindrich, Lovis, et al.
Veröffentlicht: (2025)
Efficient Global Optimization of Two-Layer ReLU Networks: Quadratic-Time Algorithms and Adversarial Training
von: Bai, Yatong, et al.
Veröffentlicht: (2022)
von: Bai, Yatong, et al.
Veröffentlicht: (2022)
Step-Level Sparse Autoencoder for Reasoning Process Interpretation
von: Yang, Xuan, et al.
Veröffentlicht: (2026)
von: Yang, Xuan, et al.
Veröffentlicht: (2026)
Low-Rank Adapting Models for Sparse Autoencoders
von: Chen, Matthew, et al.
Veröffentlicht: (2025)
von: Chen, Matthew, et al.
Veröffentlicht: (2025)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
von: Huang, Yixiao, et al.
Veröffentlicht: (2025)
von: Huang, Yixiao, et al.
Veröffentlicht: (2025)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
von: Makelov, Aleksandar, et al.
Veröffentlicht: (2024)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
von: Simon, Elana, et al.
Veröffentlicht: (2026)
von: Simon, Elana, et al.
Veröffentlicht: (2026)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
von: Marks, Luke, et al.
Veröffentlicht: (2024)
von: Marks, Luke, et al.
Veröffentlicht: (2024)
Sparse Autoencoders Trained on the Same Data Learn Different Features
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
von: Ayonrinde, Kola
Veröffentlicht: (2024)
von: Ayonrinde, Kola
Veröffentlicht: (2024)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
von: Korznikov, Anton, et al.
Veröffentlicht: (2026)
von: Korznikov, Anton, et al.
Veröffentlicht: (2026)
Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning
von: Li, Jingqi, et al.
Veröffentlicht: (2022)
von: Li, Jingqi, et al.
Veröffentlicht: (2022)
Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models
von: K, Prasanth K
Veröffentlicht: (2026)
von: K, Prasanth K
Veröffentlicht: (2026)
Identifying Reasons for Contraceptive Switching from Real-World Data Using Large Language Models
von: Miao, Brenda Y., et al.
Veröffentlicht: (2024)
von: Miao, Brenda Y., et al.
Veröffentlicht: (2024)
MoRFI: Monotonic Sparse Autoencoder Feature Identification
von: Dimakopoulos, Dimitris, et al.
Veröffentlicht: (2026)
von: Dimakopoulos, Dimitris, et al.
Veröffentlicht: (2026)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
von: Pfrommer, Samuel, et al.
Veröffentlicht: (2025) -
Revealing Treatment Non-Adherence Bias in Clinical Machine Learning Using Large Language Models
von: Liang, Zhongyuan, et al.
Veröffentlicht: (2025) -
Reinforcement Learning for Flow-Matching Policies
von: Pfrommer, Samuel, et al.
Veröffentlicht: (2025) -
Absence of spurious solutions far from ground truth: A low-rank analysis with high-order losses
von: Ma, Ziye, et al.
Veröffentlicht: (2024) -
Transport of Algebraic Structure to Latent Embeddings
von: Pfrommer, Samuel, et al.
Veröffentlicht: (2024)