Evaluating Synthetic Activations composed of SAE Latents in GPT-2
Fuente:
arXiv
Salvato in:
| Autori principali: | Giglemiani, Giorgi, Petrova, Nora, Mangat, Chatrik Singh, Janiak, Jett, Heimersheim, Stefan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Characterizing stable regions in the residual stream of LLMs
di: Janiak, Jett, et al.
Pubblicazione: (2024)
di: Janiak, Jett, et al.
Pubblicazione: (2024)
You can remove GPT2's LayerNorm by fine-tuning
di: Heimersheim, Stefan
Pubblicazione: (2024)
di: Heimersheim, Stefan
Pubblicazione: (2024)
Evolution of SAE Features Across Layers in LLMs
di: Balcells, Daniel, et al.
Pubblicazione: (2024)
di: Balcells, Daniel, et al.
Pubblicazione: (2024)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
di: Lee, Daniel J., et al.
Pubblicazione: (2024)
di: Lee, Daniel J., et al.
Pubblicazione: (2024)
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
di: Fillingham, Sean P., et al.
Pubblicazione: (2025)
di: Fillingham, Sean P., et al.
Pubblicazione: (2025)
Evaluating SAE interpretability without explanations
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
di: Janiak, Jett, et al.
Pubblicazione: (2023)
di: Janiak, Jett, et al.
Pubblicazione: (2023)
How to use and interpret activation patching
di: Heimersheim, Stefan, et al.
Pubblicazione: (2024)
di: Heimersheim, Stefan, et al.
Pubblicazione: (2024)
Boundary Point Jailbreaking of Black-Box LLMs
di: Davies, Xander, et al.
Pubblicazione: (2026)
di: Davies, Xander, et al.
Pubblicazione: (2026)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
di: Baroni, Luca, et al.
Pubblicazione: (2025)
di: Baroni, Luca, et al.
Pubblicazione: (2025)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
di: Arcuschin, Iván, et al.
Pubblicazione: (2025)
di: Arcuschin, Iván, et al.
Pubblicazione: (2025)
Latent Adversarial Training Improves the Representation of Refusal
di: Abbas, Alexandra, et al.
Pubblicazione: (2025)
di: Abbas, Alexandra, et al.
Pubblicazione: (2025)
Dense SAE Latents Are Features, Not Bugs
di: Sun, Xiaoqing, et al.
Pubblicazione: (2025)
di: Sun, Xiaoqing, et al.
Pubblicazione: (2025)
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
di: Recchia, Gabriel, et al.
Pubblicazione: (2025)
di: Recchia, Gabriel, et al.
Pubblicazione: (2025)
Detecting Strategic Deception Using Linear Probes
di: Goldowsky-Dill, Nicholas, et al.
Pubblicazione: (2025)
di: Goldowsky-Dill, Nicholas, et al.
Pubblicazione: (2025)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2026)
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2026)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
di: Braun, Dan, et al.
Pubblicazione: (2025)
di: Braun, Dan, et al.
Pubblicazione: (2025)
From Stability to Inconsistency: A Study of Moral Preferences in LLMs
di: Jotautaite, Monika, et al.
Pubblicazione: (2025)
di: Jotautaite, Monika, et al.
Pubblicazione: (2025)
Latent Regularization in Generative Test Input Generation
di: Merabishvili, Giorgi, et al.
Pubblicazione: (2026)
di: Merabishvili, Giorgi, et al.
Pubblicazione: (2026)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
di: Ayonrinde, Kola, et al.
Pubblicazione: (2024)
di: Ayonrinde, Kola, et al.
Pubblicazione: (2024)
Benchmarking Deception Probes via Black-to-White Performance Boosts
di: Parrack, Avi, et al.
Pubblicazione: (2025)
di: Parrack, Avi, et al.
Pubblicazione: (2025)
Tokenized SAEs: Disentangling SAE Reconstructions
di: Dooms, Thomas, et al.
Pubblicazione: (2025)
di: Dooms, Thomas, et al.
Pubblicazione: (2025)
SAE: Single Architecture Ensemble Neural Networks
di: Ferianc, Martin, et al.
Pubblicazione: (2024)
di: Ferianc, Martin, et al.
Pubblicazione: (2024)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
di: Vatsal, Shubham, et al.
Pubblicazione: (2024)
di: Vatsal, Shubham, et al.
Pubblicazione: (2024)
Explainable YOLO-Based Dyslexia Detection in Synthetic Handwriting Data
di: Fink, Nora
Pubblicazione: (2025)
di: Fink, Nora
Pubblicazione: (2025)
ActivationReasoning: Logical Reasoning in Latent Activation Spaces
di: Helff, Lukas, et al.
Pubblicazione: (2025)
di: Helff, Lukas, et al.
Pubblicazione: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
di: Bushnaq, Lucius, et al.
Pubblicazione: (2024)
di: Bushnaq, Lucius, et al.
Pubblicazione: (2024)
A Geometry-Based View of Mahalanobis OOD Detection
di: Janiak, Denis, et al.
Pubblicazione: (2025)
di: Janiak, Denis, et al.
Pubblicazione: (2025)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
di: Korznikov, Anton, et al.
Pubblicazione: (2025)
di: Korznikov, Anton, et al.
Pubblicazione: (2025)
Concept-SAE: Active Causal Probing of Visual Model Behavior
di: Ding, Jianrong, et al.
Pubblicazione: (2025)
di: Ding, Jianrong, et al.
Pubblicazione: (2025)
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
di: Cao, Tue M., et al.
Pubblicazione: (2026)
di: Cao, Tue M., et al.
Pubblicazione: (2026)
Activation Differences Reveal Backdoors: A Comparison of SAE Architectures
di: Kumar, Sachin
Pubblicazione: (2026)
di: Kumar, Sachin
Pubblicazione: (2026)
Latent Process Generator Matching
di: Billera, Lukas, et al.
Pubblicazione: (2026)
di: Billera, Lukas, et al.
Pubblicazione: (2026)
AlignSAE: Concept-Aligned Sparse Autoencoders
di: Yang, Minglai, et al.
Pubblicazione: (2025)
di: Yang, Minglai, et al.
Pubblicazione: (2025)
Latent Stochastic Interpolants
di: Singh, Saurabh, et al.
Pubblicazione: (2025)
di: Singh, Saurabh, et al.
Pubblicazione: (2025)
Obfuscated Activations Bypass LLM Latent-Space Defenses
di: Bailey, Luke, et al.
Pubblicazione: (2024)
di: Bailey, Luke, et al.
Pubblicazione: (2024)
SPARLING: Learning Latent Representations with Extremely Sparse Activations
di: Gupta, Kavi, et al.
Pubblicazione: (2023)
di: Gupta, Kavi, et al.
Pubblicazione: (2023)
Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
di: Janiak, Denis, et al.
Pubblicazione: (2025)
di: Janiak, Denis, et al.
Pubblicazione: (2025)
Improving the Generation and Evaluation of Synthetic Data for Downstream Medical Causal Inference
di: Amad, Harry, et al.
Pubblicazione: (2025)
di: Amad, Harry, et al.
Pubblicazione: (2025)
Music2Latent: Consistency Autoencoders for Latent Audio Compression
di: Pasini, Marco, et al.
Pubblicazione: (2024)
di: Pasini, Marco, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Characterizing stable regions in the residual stream of LLMs
di: Janiak, Jett, et al.
Pubblicazione: (2024) -
You can remove GPT2's LayerNorm by fine-tuning
di: Heimersheim, Stefan
Pubblicazione: (2024) -
Evolution of SAE Features Across Layers in LLMs
di: Balcells, Daniel, et al.
Pubblicazione: (2024) -
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
di: Lee, Daniel J., et al.
Pubblicazione: (2024) -
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
di: Fillingham, Sean P., et al.
Pubblicazione: (2025)