Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
Fuente:
arXiv
Salvato in:
| Autori principali: | Weng, Jiaqi, Zheng, Han, Zhang, Hanyu, Zhou, Ej, He, Qinqin, Tao, Jialing, Xue, Hui, Chu, Zhixuan, Wang, Xiting |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Single Neuron Works: Precise Concept Erasure in Text-to-Image Diffusion Models
di: He, Qinqin, et al.
Pubblicazione: (2025)
di: He, Qinqin, et al.
Pubblicazione: (2025)
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
di: Liang, Shuang, et al.
Pubblicazione: (2025)
di: Liang, Shuang, et al.
Pubblicazione: (2025)
Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models
di: Wang, Shuqiang, et al.
Pubblicazione: (2026)
di: Wang, Shuqiang, et al.
Pubblicazione: (2026)
Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models
di: Liang, Shuang, et al.
Pubblicazione: (2025)
di: Liang, Shuang, et al.
Pubblicazione: (2025)
Route Sparse Autoencoder to Interpret Large Language Models
di: Shi, Wei, et al.
Pubblicazione: (2025)
di: Shi, Wei, et al.
Pubblicazione: (2025)
SAIL: Sound Abstract Interpreters with LLMs
di: Gu, Qiuhan, et al.
Pubblicazione: (2026)
di: Gu, Qiuhan, et al.
Pubblicazione: (2026)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025)
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
di: Makelov, Aleksandar, et al.
Pubblicazione: (2024)
di: Makelov, Aleksandar, et al.
Pubblicazione: (2024)
Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
di: Parsan, Nithin, et al.
Pubblicazione: (2025)
di: Parsan, Nithin, et al.
Pubblicazione: (2025)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
di: Pan, Licheng, et al.
Pubblicazione: (2026)
di: Pan, Licheng, et al.
Pubblicazione: (2026)
Causal Interpretation of Sparse Autoencoder Features in Vision
di: Han, Sangyu, et al.
Pubblicazione: (2025)
di: Han, Sangyu, et al.
Pubblicazione: (2025)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
di: Yang, Junxiao, et al.
Pubblicazione: (2026)
di: Yang, Junxiao, et al.
Pubblicazione: (2026)
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
di: Morelli, Fabian, et al.
Pubblicazione: (2026)
di: Morelli, Fabian, et al.
Pubblicazione: (2026)
Interpretable LLM Guardrails via Sparse Representation Steering
di: He, Zeqing, et al.
Pubblicazione: (2025)
di: He, Zeqing, et al.
Pubblicazione: (2025)
LFRAG: Layout-oriented Fine-grained Retrieval-Augmented Generation on Multimodal Document Understanding
di: Zhu, Yifan, et al.
Pubblicazione: (2026)
di: Zhu, Yifan, et al.
Pubblicazione: (2026)
SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
di: Deng, Junyuan, et al.
Pubblicazione: (2025)
di: Deng, Junyuan, et al.
Pubblicazione: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
Interpretable Company Similarity with Sparse Autoencoders
di: Molinari, Marco, et al.
Pubblicazione: (2024)
di: Molinari, Marco, et al.
Pubblicazione: (2024)
Interpreting CLIP with Hierarchical Sparse Autoencoders
di: Zaigrajew, Vladimir, et al.
Pubblicazione: (2025)
di: Zaigrajew, Vladimir, et al.
Pubblicazione: (2025)
Enhancing Safety of Large Language Models via Embedding Space Separation
di: Zhao, Xu, et al.
Pubblicazione: (2026)
di: Zhao, Xu, et al.
Pubblicazione: (2026)
Bias Beyond English: Evaluating Social Bias and Debiasing Methods in a Low-Resource Setting
di: Zhou, Ej, et al.
Pubblicazione: (2025)
di: Zhou, Ej, et al.
Pubblicazione: (2025)
Navigating the Landscape of Large Language Models: A Comprehensive Review and Analysis of Paradigms and Fine-Tuning Strategies
di: Weng, Benjue
Pubblicazione: (2024)
di: Weng, Benjue
Pubblicazione: (2024)
Controlling Large Language Models Through Concept Activation Vectors
di: Zhang, Hanyu, et al.
Pubblicazione: (2025)
di: Zhang, Hanyu, et al.
Pubblicazione: (2025)
Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
di: Zhao, Daniel, et al.
Pubblicazione: (2025)
di: Zhao, Daniel, et al.
Pubblicazione: (2025)
Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information
di: Wang, Shih-Heng, et al.
Pubblicazione: (2026)
di: Wang, Shih-Heng, et al.
Pubblicazione: (2026)
SAIL: Unsupervised Spatial-Angular Interpretable Feature Learning for RF Map Synthesis
di: Sarkar, Sopan, et al.
Pubblicazione: (2026)
di: Sarkar, Sopan, et al.
Pubblicazione: (2026)
Interpreting CFD Surrogates through Sparse Autoencoders
di: Hu, Yeping, et al.
Pubblicazione: (2025)
di: Hu, Yeping, et al.
Pubblicazione: (2025)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
di: Tolooshams, Bahareh, et al.
Pubblicazione: (2025)
di: Tolooshams, Bahareh, et al.
Pubblicazione: (2025)
Interpretable Reward Model via Sparse Autoencoder
di: Zhang, Shuyi, et al.
Pubblicazione: (2025)
di: Zhang, Shuyi, et al.
Pubblicazione: (2025)
Interpreting Attention Layer Outputs with Sparse Autoencoders
di: Kissane, Connor, et al.
Pubblicazione: (2024)
di: Kissane, Connor, et al.
Pubblicazione: (2024)
Toward Identifiable Sparse Autoencoders
di: Nelson, Walter, et al.
Pubblicazione: (2026)
di: Nelson, Walter, et al.
Pubblicazione: (2026)
SAIL-VL2 Technical Report
di: Yin, Weijie, et al.
Pubblicazione: (2025)
di: Yin, Weijie, et al.
Pubblicazione: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Reinfier and Reintrainer: Verification and Interpretation-Driven Safe Deep Reinforcement Learning Frameworks
di: Yang, Zixuan, et al.
Pubblicazione: (2024)
di: Yang, Zixuan, et al.
Pubblicazione: (2024)
CE-Bench: Towards a Reliable Contrastive Evaluation Benchmark of Interpretability of Sparse Autoencoders
di: Gulko, Alex, et al.
Pubblicazione: (2025)
di: Gulko, Alex, et al.
Pubblicazione: (2025)
Modular SAIL: dream or reality?
di: Kourzanov, Petr, et al.
Pubblicazione: (2025)
di: Kourzanov, Petr, et al.
Pubblicazione: (2025)
LET JAPAN SAIL FORTH
Pubblicazione: (1999)
Pubblicazione: (1999)
SAIL Thomson Reuters Update
di: Culp, Kristin
Pubblicazione: ()
di: Culp, Kristin
Pubblicazione: ()
Towards Unified Facial Action Unit Recognition Framework by Large Language Models
di: Hu, Guohong, et al.
Pubblicazione: (2024)
di: Hu, Guohong, et al.
Pubblicazione: (2024)
Vision Language Model for Interpretable and Fine-grained Detection of Safety Compliance in Diverse Workplaces
di: Chen, Zhiling, et al.
Pubblicazione: (2024)
di: Chen, Zhiling, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Single Neuron Works: Precise Concept Erasure in Text-to-Image Diffusion Models
di: He, Qinqin, et al.
Pubblicazione: (2025) -
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
di: Liang, Shuang, et al.
Pubblicazione: (2025) -
Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models
di: Wang, Shuqiang, et al.
Pubblicazione: (2026) -
Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models
di: Liang, Shuang, et al.
Pubblicazione: (2025) -
Route Sparse Autoencoder to Interpret Large Language Models
di: Shi, Wei, et al.
Pubblicazione: (2025)