Gespeichert in:
| Hauptverfasser: | Janiak, Jett, Karwowski, Jacek, Mangat, Chatrik Singh, Giglemiani, Giorgi, Petrova, Nora, Heimersheim, Stefan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.17113 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
Boundary Point Jailbreaking of Black-Box LLMs
von: Davies, Xander, et al.
Veröffentlicht: (2026)
von: Davies, Xander, et al.
Veröffentlicht: (2026)
Incoherence in Goal-Conditioned Autoregressive Models
von: Karwowski, Jacek, et al.
Veröffentlicht: (2025)
von: Karwowski, Jacek, et al.
Veröffentlicht: (2025)
Hilbert geometry of the symmetric positive-definite bicone: Application to the geometry of the extended Gaussian family
von: Karwowski, Jacek, et al.
Veröffentlicht: (2025)
von: Karwowski, Jacek, et al.
Veröffentlicht: (2025)
Geometric structures and deviations on James' symmetric positive-definite matrix bicone domain
von: Karwowski, Jacek, et al.
Veröffentlicht: (2026)
von: Karwowski, Jacek, et al.
Veröffentlicht: (2026)
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
von: Janiak, Jett, et al.
Veröffentlicht: (2023)
von: Janiak, Jett, et al.
Veröffentlicht: (2023)
You can remove GPT2's LayerNorm by fine-tuning
von: Heimersheim, Stefan
Veröffentlicht: (2024)
von: Heimersheim, Stefan
Veröffentlicht: (2024)
From Stability to Inconsistency: A Study of Moral Preferences in LLMs
von: Jotautaite, Monika, et al.
Veröffentlicht: (2025)
von: Jotautaite, Monika, et al.
Veröffentlicht: (2025)
How to use and interpret activation patching
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
Evolution of SAE Features Across Layers in LLMs
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
von: Recchia, Gabriel, et al.
Veröffentlicht: (2025)
von: Recchia, Gabriel, et al.
Veröffentlicht: (2025)
Likelihood hacking in probabilistic program synthesis
von: Karwowski, Jacek, et al.
Veröffentlicht: (2026)
von: Karwowski, Jacek, et al.
Veröffentlicht: (2026)
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025)
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025)
Detecting Strategic Deception Using Linear Probes
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
von: Braun, Dan, et al.
Veröffentlicht: (2025)
von: Braun, Dan, et al.
Veröffentlicht: (2025)
Benchmarking Deception Probes via Black-to-White Performance Boosts
von: Parrack, Avi, et al.
Veröffentlicht: (2025)
von: Parrack, Avi, et al.
Veröffentlicht: (2025)
Transformers represent belief state geometry in their residual stream
von: Shai, Adam S., et al.
Veröffentlicht: (2024)
von: Shai, Adam S., et al.
Veröffentlicht: (2024)
Hallucination Detection in LLMs Using Spectral Features of Attention Maps
von: Binkowski, Jakub, et al.
Veröffentlicht: (2025)
von: Binkowski, Jakub, et al.
Veröffentlicht: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
Latent Adversarial Training Improves the Representation of Refusal
von: Abbas, Alexandra, et al.
Veröffentlicht: (2025)
von: Abbas, Alexandra, et al.
Veröffentlicht: (2025)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
A Geometry-Based View of Mahalanobis OOD Detection
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
Erfonium: A Hooke Atom with Soft Interaction Potential
von: Karwowski, Jacek, et al.
Veröffentlicht: (2023)
von: Karwowski, Jacek, et al.
Veröffentlicht: (2023)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Confirmation bias: A challenge for scalable oversight
von: Recchia, Gabriel, et al.
Veröffentlicht: (2025)
von: Recchia, Gabriel, et al.
Veröffentlicht: (2025)
Refusal in LLMs is an Affine Function
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Thermal Robustness of Retrieval in Dense Associative Memories: LSE vs LSR Kernels
von: Petrova, Tatiana
Veröffentlicht: (2026)
von: Petrova, Tatiana
Veröffentlicht: (2026)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
von: Kowal, Matthew, et al.
Veröffentlicht: (2026)
Sycophantic Anchors: Localizing and Quantifying User Agreement in Reasoning Models
von: Duszenko, Jacek
Veröffentlicht: (2026)
von: Duszenko, Jacek
Veröffentlicht: (2026)
COMBINEX: A Unified Counterfactual Explainer for Graph Neural Networks via Node Feature and Structural Perturbations
von: Giorgi, Flavio, et al.
Veröffentlicht: (2025)
von: Giorgi, Flavio, et al.
Veröffentlicht: (2025)
Bootstrap Sampling Rate Greater than 1.0 May Improve Random Forest Performance
von: Kaźmierczak, Stanisław, et al.
Veröffentlicht: (2024)
von: Kaźmierczak, Stanisław, et al.
Veröffentlicht: (2024)
Interpretable Multi-task Learning with Shared Variable Embeddings
von: Żelaszczyk, Maciej, et al.
Veröffentlicht: (2024)
von: Żelaszczyk, Maciej, et al.
Veröffentlicht: (2024)
A-PETE: Adaptive Prototype Explanations of Tree Ensembles
von: Karolczak, Jacek, et al.
Veröffentlicht: (2024)
von: Karolczak, Jacek, et al.
Veröffentlicht: (2024)
Towards Unbiased Calibration using Meta-Regularization
von: Wang, Cheng, et al.
Veröffentlicht: (2023)
von: Wang, Cheng, et al.
Veröffentlicht: (2023)
Alike Parts: A Feature-Informed Approach to Local and Global Prototype Explanations
von: Karolczak, Jacek, et al.
Veröffentlicht: (2026)
von: Karolczak, Jacek, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024) -
Boundary Point Jailbreaking of Black-Box LLMs
von: Davies, Xander, et al.
Veröffentlicht: (2026) -
Incoherence in Goal-Conditioned Autoregressive Models
von: Karwowski, Jacek, et al.
Veröffentlicht: (2025) -
Hilbert geometry of the symmetric positive-definite bicone: Application to the geometry of the extended Gaussian family
von: Karwowski, Jacek, et al.
Veröffentlicht: (2025) -
Geometric structures and deviations on James' symmetric positive-definite matrix bicone domain
von: Karwowski, Jacek, et al.
Veröffentlicht: (2026)