Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Aaron J., Srinivas, Suraj, Bhalla, Usha, Lakkaraju, Himabindu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2023)
di: Bhalla, Usha, et al.
Pubblicazione: (2023)
Certifying LLM Safety against Adversarial Prompting
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
Towards Unifying Interpretability and Control: Evaluation via Intervention
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
di: Oesterling, Alex, et al.
Pubblicazione: (2024)
di: Oesterling, Alex, et al.
Pubblicazione: (2024)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
di: Badrinath, Charumathi, et al.
Pubblicazione: (2024)
di: Badrinath, Charumathi, et al.
Pubblicazione: (2024)
Towards Interpretable Soft Prompts
di: Patel, Oam, et al.
Pubblicazione: (2025)
di: Patel, Oam, et al.
Pubblicazione: (2025)
Generalized Group Data Attribution
di: Ley, Dan, et al.
Pubblicazione: (2024)
di: Ley, Dan, et al.
Pubblicazione: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
Learning Recourse Costs from Pairwise Feature Comparisons
di: Rawal, Kaivalya, et al.
Pubblicazione: (2024)
di: Rawal, Kaivalya, et al.
Pubblicazione: (2024)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
di: Li, Aaron J., et al.
Pubblicazione: (2024)
di: Li, Aaron J., et al.
Pubblicazione: (2024)
Characterizing Data Point Vulnerability via Average-Case Robustness
di: Han, Tessa, et al.
Pubblicazione: (2023)
di: Han, Tessa, et al.
Pubblicazione: (2023)
Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
di: Peng, Kenny, et al.
Pubblicazione: (2025)
di: Peng, Kenny, et al.
Pubblicazione: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
Do Sparse Autoencoders Capture Concept Manifolds?
di: Bhalla, Usha, et al.
Pubblicazione: (2026)
di: Bhalla, Usha, et al.
Pubblicazione: (2026)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
di: Joshi, Shruti, et al.
Pubblicazione: (2025)
di: Joshi, Shruti, et al.
Pubblicazione: (2025)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
di: Du, Hongzhe, et al.
Pubblicazione: (2025)
di: Du, Hongzhe, et al.
Pubblicazione: (2025)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
di: Muhamed, Aashiq, et al.
Pubblicazione: (2024)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2024)
Towards Understanding the Robustness of Sparse Autoencoders
di: Saiyed, Ahson, et al.
Pubblicazione: (2026)
di: Saiyed, Ahson, et al.
Pubblicazione: (2026)
A Study on the Calibration of In-context Learning
di: Zhang, Hanlin, et al.
Pubblicazione: (2023)
di: Zhang, Hanlin, et al.
Pubblicazione: (2023)
Manipulating Large Language Models to Increase Product Visibility
di: Kumar, Aounon, et al.
Pubblicazione: (2024)
di: Kumar, Aounon, et al.
Pubblicazione: (2024)
EvoLM: In Search of Lost Language Model Training Dynamics
di: Qi, Zhenting, et al.
Pubblicazione: (2025)
di: Qi, Zhenting, et al.
Pubblicazione: (2025)
Sparse Autoencoder Features for Classifications and Transferability
di: Gallifant, Jack, et al.
Pubblicazione: (2025)
di: Gallifant, Jack, et al.
Pubblicazione: (2025)
How Much Can We Forget about Data Contamination?
di: Bordt, Sebastian, et al.
Pubblicazione: (2024)
di: Bordt, Sebastian, et al.
Pubblicazione: (2024)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
di: Sainsbury, Chris, et al.
Pubblicazione: (2026)
di: Sainsbury, Chris, et al.
Pubblicazione: (2026)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
di: Muchane, Mark, et al.
Pubblicazione: (2025)
di: Muchane, Mark, et al.
Pubblicazione: (2025)
Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
di: Farnik, Lucy, et al.
Pubblicazione: (2025)
di: Farnik, Lucy, et al.
Pubblicazione: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
di: Xiong, Zidi, et al.
Pubblicazione: (2026)
di: Xiong, Zidi, et al.
Pubblicazione: (2026)
Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
di: Li, T. Ed, et al.
Pubblicazione: (2025)
di: Li, T. Ed, et al.
Pubblicazione: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
di: Zhu, Xudong, et al.
Pubblicazione: (2025)
di: Zhu, Xudong, et al.
Pubblicazione: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
di: Jing, Yi, et al.
Pubblicazione: (2026)
di: Jing, Yi, et al.
Pubblicazione: (2026)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025) -
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2023) -
Certifying LLM Safety against Adversarial Prompting
di: Kumar, Aounon, et al.
Pubblicazione: (2023) -
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
di: Bhalla, Usha, et al.
Pubblicazione: (2024) -
Towards Unifying Interpretability and Control: Evaluation via Intervention
di: Bhalla, Usha, et al.
Pubblicazione: (2024)