Group Crosscoders for Mechanistic Analysis of Symmetry
Fuente:
arXiv
Saved in:
| Main Author: | Gorton, Liv |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Missing Curve Detectors of InceptionV1: Applying Sparse Autoencoders to InceptionV1 Early Vision
by: Gorton, Liv
Published: (2024)
by: Gorton, Liv
Published: (2024)
Adversarial Examples Are Not Bugs, They Are Superposition
by: Gorton, Liv, et al.
Published: (2025)
by: Gorton, Liv, et al.
Published: (2025)
Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
by: Kassem, Aly, et al.
Published: (2026)
by: Kassem, Aly, et al.
Published: (2026)
Understanding sparse autoencoder scaling in the presence of feature manifolds
by: Michaud, Eric J., et al.
Published: (2025)
by: Michaud, Eric J., et al.
Published: (2025)
Sparse Crosscoders for diffing MoEs and Dense models
by: Chaudhari, Marmik, et al.
Published: (2026)
by: Chaudhari, Marmik, et al.
Published: (2026)
fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery
by: Demou, Andreas D., et al.
Published: (2026)
by: Demou, Andreas D., et al.
Published: (2026)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
by: Minder, Julian, et al.
Published: (2025)
by: Minder, Julian, et al.
Published: (2025)
Measuring the Representational Alignment of Neural Systems in Superposition
by: Liu, Sunny, et al.
Published: (2026)
by: Liu, Sunny, et al.
Published: (2026)
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
by: Jiralerspong, Thomas, et al.
Published: (2026)
by: Jiralerspong, Thomas, et al.
Published: (2026)
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
by: Bayazit, Deniz, et al.
Published: (2025)
by: Bayazit, Deniz, et al.
Published: (2025)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
by: Erdogan, Ege, et al.
Published: (2025)
by: Erdogan, Ege, et al.
Published: (2025)
Learning Environment for the Air Domain (LEAD)
by: Strand, Andreas, et al.
Published: (2023)
by: Strand, Andreas, et al.
Published: (2023)
Root Cause Analysis of Measurement and Mechanistic Anomalies
by: Suhr, Hendrik, et al.
Published: (2026)
by: Suhr, Hendrik, et al.
Published: (2026)
Mechanistic Analysis of Circuit Preservation in Federated Learning
by: Haseeb, Muhammad, et al.
Published: (2025)
by: Haseeb, Muhammad, et al.
Published: (2025)
Non-parametric Hypothesis Tests for Distributional Group Symmetry
by: Chiu, Kenny, et al.
Published: (2023)
by: Chiu, Kenny, et al.
Published: (2023)
Symmetry From Scratch: Group Equivariance as a Supervised Learning Task
by: Huang, Haozhe, et al.
Published: (2024)
by: Huang, Haozhe, et al.
Published: (2024)
Sample Complexity of Probability Divergences under Group Symmetry
by: Chen, Ziyu, et al.
Published: (2023)
by: Chen, Ziyu, et al.
Published: (2023)
Continuous Symmetry Discovery and Enforcement Using Infinitesimal Generators of Multi-parameter Group Actions
by: Shaw, Ben, et al.
Published: (2025)
by: Shaw, Ben, et al.
Published: (2025)
Discovering Symmetry Breaking in Physical Systems with Relaxed Group Convolution
by: Wang, Rui, et al.
Published: (2023)
by: Wang, Rui, et al.
Published: (2023)
SymmPI: Predictive Inference for Data with Group Symmetries
by: Dobriban, Edgar, et al.
Published: (2023)
by: Dobriban, Edgar, et al.
Published: (2023)
A Mechanistic Analysis of Looped Reasoning Language Models
by: Blayney, Hugh, et al.
Published: (2026)
by: Blayney, Hugh, et al.
Published: (2026)
Conditional Predictive Inference for General Structured Data with Group Symmetries
by: Shen, Yichen, et al.
Published: (2026)
by: Shen, Yichen, et al.
Published: (2026)
Dictionary Learning under Symmetries via Group Representations
by: Ghosh, Subhroshekhar, et al.
Published: (2023)
by: Ghosh, Subhroshekhar, et al.
Published: (2023)
Group-Invariant Unsupervised Skill Discovery: Symmetry-aware Skill Representations for Generalizable Behavior
by: Chang, Junwoo, et al.
Published: (2026)
by: Chang, Junwoo, et al.
Published: (2026)
When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability
by: Ronge, Raphael, et al.
Published: (2026)
by: Ronge, Raphael, et al.
Published: (2026)
Probing Ranking LLMs: A Mechanistic Analysis for Information Retrieval
by: Chowdhury, Tanya, et al.
Published: (2024)
by: Chowdhury, Tanya, et al.
Published: (2024)
Exemplar Partitioning for Mechanistic Interpretability
by: Rumbelow, Jessica
Published: (2026)
by: Rumbelow, Jessica
Published: (2026)
Open Problems in Mechanistic Interpretability
by: Sharkey, Lee, et al.
Published: (2025)
by: Sharkey, Lee, et al.
Published: (2025)
From Mechanistic to Compositional Interpretability
by: Gauderis, Ward, et al.
Published: (2026)
by: Gauderis, Ward, et al.
Published: (2026)
A Mechanistic Analysis of Transformers for Dynamical Systems
by: Duthé, Gregory, et al.
Published: (2025)
by: Duthé, Gregory, et al.
Published: (2025)
WARP-LUTs -- Walsh-Assisted Relaxation for Probabilistic Look Up Tables
by: Gerlach, Lino, et al.
Published: (2025)
by: Gerlach, Lino, et al.
Published: (2025)
WARP Logic Neural Networks
by: Gerlach, Lino, et al.
Published: (2026)
by: Gerlach, Lino, et al.
Published: (2026)
A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task
by: Brinkmann, Jannik, et al.
Published: (2024)
by: Brinkmann, Jannik, et al.
Published: (2024)
Current Symmetry Group Equivariant Convolution Frameworks for Representation Learning
by: Basheer, Ramzan, et al.
Published: (2024)
by: Basheer, Ramzan, et al.
Published: (2024)
Exploring Grokking: Experimental and Mechanistic Investigations
by: Qiye, Hu, et al.
Published: (2024)
by: Qiye, Hu, et al.
Published: (2024)
Mechanistic Interpretability of Reinforcement Learning Agents
by: Trim, Tristan, et al.
Published: (2024)
by: Trim, Tristan, et al.
Published: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
by: Palumbo, Nils, et al.
Published: (2024)
by: Palumbo, Nils, et al.
Published: (2024)
Mechanistic Design and Scaling of Hybrid Architectures
by: Poli, Michael, et al.
Published: (2024)
by: Poli, Michael, et al.
Published: (2024)
Mechanistic Interpretability for Neural TSP Solvers
by: Narad, Reuben, et al.
Published: (2025)
by: Narad, Reuben, et al.
Published: (2025)
Models of Heavy-Tailed Mechanistic Universality
by: Hodgkinson, Liam, et al.
Published: (2025)
by: Hodgkinson, Liam, et al.
Published: (2025)
Similar Items
-
The Missing Curve Detectors of InceptionV1: Applying Sparse Autoencoders to InceptionV1 Early Vision
by: Gorton, Liv
Published: (2024) -
Adversarial Examples Are Not Bugs, They Are Superposition
by: Gorton, Liv, et al.
Published: (2025) -
Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
by: Kassem, Aly, et al.
Published: (2026) -
Understanding sparse autoencoder scaling in the presence of feature manifolds
by: Michaud, Eric J., et al.
Published: (2025) -
Sparse Crosscoders for diffing MoEs and Dense models
by: Chaudhari, Marmik, et al.
Published: (2026)