SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Härle, Ruben, Friedrich, Felix, Brack, Manuel, Deiseroth, Björn, Schramowski, Patrick, Kersting, Kristian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring and Guiding Monosemanticity
by: Härle, Ruben, et al.
Published: (2025)
by: Härle, Ruben, et al.
Published: (2025)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
by: Deiseroth, Björn, et al.
Published: (2024)
by: Deiseroth, Björn, et al.
Published: (2024)
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench
by: Friedrich, Felix, et al.
Published: (2025)
by: Friedrich, Felix, et al.
Published: (2025)
LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies
by: Friedrich, Felix, et al.
Published: (2024)
by: Friedrich, Felix, et al.
Published: (2024)
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You
by: Friedrich, Felix, et al.
Published: (2024)
by: Friedrich, Felix, et al.
Published: (2024)
Divergent Token Metrics: Measuring degradation to prune away LLM components -- and optimize quantization
by: Deiseroth, Björn, et al.
Published: (2023)
by: Deiseroth, Björn, et al.
Published: (2023)
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
by: Deiseroth, Björn, et al.
Published: (2023)
by: Deiseroth, Björn, et al.
Published: (2023)
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
by: Helff, Lukas, et al.
Published: (2024)
by: Helff, Lukas, et al.
Published: (2024)
Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
by: Struppek, Lukas, et al.
Published: (2022)
by: Struppek, Lukas, et al.
Published: (2022)
CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models
by: Hegde, Niharika, et al.
Published: (2025)
by: Hegde, Niharika, et al.
Published: (2025)
Does CLIP Know My Face?
by: Hintersdorf, Dominik, et al.
Published: (2022)
by: Hintersdorf, Dominik, et al.
Published: (2022)
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
by: Ghussin, Yusser Al, et al.
Published: (2026)
by: Ghussin, Yusser Al, et al.
Published: (2026)
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
by: Höth, Max Henning, et al.
Published: (2026)
by: Höth, Max Henning, et al.
Published: (2026)
Core Tokensets for Data-efficient Sequential Training of Transformers
by: Paul, Subarnaduti, et al.
Published: (2024)
by: Paul, Subarnaduti, et al.
Published: (2024)
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
by: Deiseroth, Björn, et al.
Published: (2025)
by: Deiseroth, Björn, et al.
Published: (2025)
ActivationReasoning: Logical Reasoning in Latent Activation Spaces
by: Helff, Lukas, et al.
Published: (2025)
by: Helff, Lukas, et al.
Published: (2025)
LEDITS++: Limitless Image Editing using Text-to-Image Models
by: Brack, Manuel, et al.
Published: (2023)
by: Brack, Manuel, et al.
Published: (2023)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
by: Brack, Manuel, et al.
Published: (2025)
by: Brack, Manuel, et al.
Published: (2025)
ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
by: Tedeschi, Simone, et al.
Published: (2024)
by: Tedeschi, Simone, et al.
Published: (2024)
A Typology for Exploring the Mitigation of Shortcut Behavior
by: Friedrich, Felix, et al.
Published: (2022)
by: Friedrich, Felix, et al.
Published: (2022)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
by: Helff, Lukas, et al.
Published: (2026)
by: Helff, Lukas, et al.
Published: (2026)
DeiSAM: Segment Anything with Deictic Prompting
by: Shindo, Hikaru, et al.
Published: (2024)
by: Shindo, Hikaru, et al.
Published: (2024)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
No Safe Dose: How Training Data Drives Unsafe Image Generation
by: Friedrich, Felix, et al.
Published: (2026)
by: Friedrich, Felix, et al.
Published: (2026)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
by: Wu, Xuansheng, et al.
Published: (2025)
by: Wu, Xuansheng, et al.
Published: (2025)
SLR: Automated Synthesis for Scalable Logical Reasoning
by: Helff, Lukas, et al.
Published: (2025)
by: Helff, Lukas, et al.
Published: (2025)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
by: Jørgensen, Mikkel Godsk, et al.
Published: (2026)
by: Jørgensen, Mikkel Godsk, et al.
Published: (2026)
SteerRM: Debiasing Reward Models via Sparse Autoencoders
by: Sun, Mengyuan, et al.
Published: (2026)
by: Sun, Mengyuan, et al.
Published: (2026)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
by: Zhang, Ruikang, et al.
Published: (2026)
by: Zhang, Ruikang, et al.
Published: (2026)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
by: Fang, Yi, et al.
Published: (2026)
by: Fang, Yi, et al.
Published: (2026)
Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders
by: Kuznetsov, Kristian, et al.
Published: (2025)
by: Kuznetsov, Kristian, et al.
Published: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
by: Chalnev, Sviatoslav, et al.
Published: (2024)
by: Chalnev, Sviatoslav, et al.
Published: (2024)
AlignSAE: Concept-Aligned Sparse Autoencoders
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
A Comparative Analysis of Sparse Autoencoder and Activation Difference in Language Model Steering
by: Xie, Jiaqing
Published: (2025)
by: Xie, Jiaqing
Published: (2025)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
by: Hua, Zhenglin, et al.
Published: (2025)
by: Hua, Zhenglin, et al.
Published: (2025)
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
by: Neitemeier, Pit, et al.
Published: (2025)
by: Neitemeier, Pit, et al.
Published: (2025)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
by: Yu, Zhuohao, et al.
Published: (2025)
by: Yu, Zhuohao, et al.
Published: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
Similar Items
-
Measuring and Guiding Monosemanticity
by: Härle, Ruben, et al.
Published: (2025) -
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
by: Deiseroth, Björn, et al.
Published: (2024) -
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
by: Sztwiertnia, Sebastian, et al.
Published: (2025) -
Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench
by: Friedrich, Felix, et al.
Published: (2025) -
LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies
by: Friedrich, Felix, et al.
Published: (2024)