Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pacela, Vitória Barin, Joshi, Shruti, Camacho, Isabela, Lacoste-Julien, Simon, Klindt, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Operationalizing Quantized Disentanglement
von: Barin-Pacela, Vitoria, et al.
Veröffentlicht: (2025)
von: Barin-Pacela, Vitoria, et al.
Veröffentlicht: (2025)
On the Identifiability of Quantized Factors
von: Barin-Pacela, Vitória, et al.
Veröffentlicht: (2023)
von: Barin-Pacela, Vitória, et al.
Veröffentlicht: (2023)
Causality is Key for Interpretability Claims to Generalise
von: Joshi, Shruti, et al.
Veröffentlicht: (2026)
von: Joshi, Shruti, et al.
Veröffentlicht: (2026)
Data Whitening Improves Sparse Autoencoder Learning
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
von: Joshi, Shruti, et al.
Veröffentlicht: (2025)
von: Joshi, Shruti, et al.
Veröffentlicht: (2025)
Probing the Representational Power of Sparse Autoencoders in Vision Models
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2025)
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2025)
Dual Optimistic Ascent (PI Control) is the Augmented Lagrangian Method in Disguise
von: Ramirez, Juan, et al.
Veröffentlicht: (2025)
von: Ramirez, Juan, et al.
Veröffentlicht: (2025)
The Impact of Off-Policy Training Data on Probe Generalisation
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
Nonparametric Partial Disentanglement via Mechanism Sparsity: Sparse Actions, Interventions and Sparse Temporal Dependencies
von: Lachapelle, Sébastien, et al.
Veröffentlicht: (2024)
von: Lachapelle, Sébastien, et al.
Veröffentlicht: (2024)
Posterior-Calibrated Causal Circuits in Variational Autoencoders: Why Image-Domain Interpretability Fails on Tabular Data
von: Roy, Dip, et al.
Veröffentlicht: (2026)
von: Roy, Dip, et al.
Veröffentlicht: (2026)
Balancing Act: Constraining Disparate Impact in Sparse Models
von: Hashemizadeh, Meraj, et al.
Veröffentlicht: (2023)
von: Hashemizadeh, Meraj, et al.
Veröffentlicht: (2023)
Position: Adopt Constraints Over Fixed Penalties in Deep Learning
von: Ramirez, Juan, et al.
Veröffentlicht: (2025)
von: Ramirez, Juan, et al.
Veröffentlicht: (2025)
Substance Beats Style: Why Beginning Students Fail to Code with LLMs
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
Superposition disentanglement of neural representations reveals hidden alignment
von: Longon, André, et al.
Veröffentlicht: (2025)
von: Longon, André, et al.
Veröffentlicht: (2025)
Empirical Evaluation of Progressive Coding for Sparse Autoencoders
von: Peter, Hans, et al.
Veröffentlicht: (2025)
von: Peter, Hans, et al.
Veröffentlicht: (2025)
Detecting Strategic Deception Using Linear Probes
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
Weight-Sharing Regularization
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2023)
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2023)
Are Sparse Autoencoder Benchmarks Reliable?
von: Chanin, David
Veröffentlicht: (2026)
von: Chanin, David
Veröffentlicht: (2026)
Sparse Autoencoders, Again?
von: Lu, Yin, et al.
Veröffentlicht: (2025)
von: Lu, Yin, et al.
Veröffentlicht: (2025)
Occam's Razor for Self Supervised Learning: What is Sufficient to Learn Good Representations?
von: Ibrahim, Mark, et al.
Veröffentlicht: (2024)
von: Ibrahim, Mark, et al.
Veröffentlicht: (2024)
On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
von: Simon, Elana, et al.
Veröffentlicht: (2026)
von: Simon, Elana, et al.
Veröffentlicht: (2026)
Ensembling Sparse Autoencoders
von: Gadgil, Soham, et al.
Veröffentlicht: (2025)
von: Gadgil, Soham, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
What Linear Probes Miss: Multi-View Probing for Weight-Space Learning
von: Heo, Eunwoo, et al.
Veröffentlicht: (2026)
von: Heo, Eunwoo, et al.
Veröffentlicht: (2026)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
von: Oldfield, James, et al.
Veröffentlicht: (2025)
von: Oldfield, James, et al.
Veröffentlicht: (2025)
Why Safety Probes Catch Liars But Miss Fanatics
von: Haralambiev, Kristiyan
Veröffentlicht: (2026)
von: Haralambiev, Kristiyan
Veröffentlicht: (2026)
Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage
von: Milligan, Alan, et al.
Veröffentlicht: (2026)
von: Milligan, Alan, et al.
Veröffentlicht: (2026)
On Importance of Code-Mixed Embeddings for Hate Speech Identification
von: Jagdale, Shruti, et al.
Veröffentlicht: (2024)
von: Jagdale, Shruti, et al.
Veröffentlicht: (2024)
Toward Identifiable Sparse Autoencoders
von: Nelson, Walter, et al.
Veröffentlicht: (2026)
von: Nelson, Walter, et al.
Veröffentlicht: (2026)
Analysis of Variational Sparse Autoencoders
von: Baker, Zachary, et al.
Veröffentlicht: (2025)
von: Baker, Zachary, et al.
Veröffentlicht: (2025)
Why Domain Generalization Fail? A View of Necessity and Sufficiency
von: Vuong, Long-Tung, et al.
Veröffentlicht: (2025)
von: Vuong, Long-Tung, et al.
Veröffentlicht: (2025)
Why Does Agentic Safety Fail to Generalize Across Tasks?
von: Slutzky, Yonatan, et al.
Veröffentlicht: (2026)
von: Slutzky, Yonatan, et al.
Veröffentlicht: (2026)
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
von: Njaradi, Valentina, et al.
Veröffentlicht: (2026)
von: Njaradi, Valentina, et al.
Veröffentlicht: (2026)
Cooper: A Library for Constrained Optimization in Deep Learning
von: Gallego-Posada, Jose, et al.
Veröffentlicht: (2025)
von: Gallego-Posada, Jose, et al.
Veröffentlicht: (2025)
In-Context Compositional Learning via Sparse Coding Transformer
von: Chen, Wei, et al.
Veröffentlicht: (2025)
von: Chen, Wei, et al.
Veröffentlicht: (2025)
Steering Language Model Refusal with Sparse Autoencoders
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
Convex SGD: Generalization Without Early Stopping
von: Hendrickx, Julien, et al.
Veröffentlicht: (2024)
von: Hendrickx, Julien, et al.
Veröffentlicht: (2024)
Can Linear Probes Measure LLM Uncertainty?
von: Dakhmouche, Ramzi, et al.
Veröffentlicht: (2025)
von: Dakhmouche, Ramzi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Operationalizing Quantized Disentanglement
von: Barin-Pacela, Vitoria, et al.
Veröffentlicht: (2025) -
On the Identifiability of Quantized Factors
von: Barin-Pacela, Vitória, et al.
Veröffentlicht: (2023) -
Causality is Key for Interpretability Claims to Generalise
von: Joshi, Shruti, et al.
Veröffentlicht: (2026) -
Data Whitening Improves Sparse Autoencoder Learning
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025) -
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2024)