Saved in:
| Main Authors: | Frikha, Ahmed, Razi, Muhammad Reza Ar, Nakka, Krishna Kanth, Mendes, Ricardo, Jiang, Xue, Zhou, Xuebing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.11232 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IncogniText: Privacy-enhancing Conditional Text Anonymization via LLM-based Private Attribute Randomization
by: Frikha, Ahmed, et al.
Published: (2024)
by: Frikha, Ahmed, et al.
Published: (2024)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
by: Nakka, Krishna Kanth, et al.
Published: (2024)
by: Nakka, Krishna Kanth, et al.
Published: (2024)
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
by: Nakka, Krishna Kanth, et al.
Published: (2024)
by: Nakka, Krishna Kanth, et al.
Published: (2024)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
by: Frikha, Ahmed, et al.
Published: (2024)
by: Frikha, Ahmed, et al.
Published: (2024)
Mammo-SAE: Interpreting Breast Cancer Concept Learning with Sparse Autoencoders
by: Nakka, Krishna Kanth
Published: (2025)
by: Nakka, Krishna Kanth
Published: (2025)
PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage
by: Nakka, Krishna Kanth, et al.
Published: (2025)
by: Nakka, Krishna Kanth, et al.
Published: (2025)
From Steering to Pedalling: Do Autonomous Driving VLMs Generalize to Cyclist-Assistive Spatial Perception and Planning?
by: Nakka, Krishna Kanth, et al.
Published: (2026)
by: Nakka, Krishna Kanth, et al.
Published: (2026)
NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
by: Nakka, Krishna Kanth, et al.
Published: (2025)
by: Nakka, Krishna Kanth, et al.
Published: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
by: Marks, Luke, et al.
Published: (2024)
by: Marks, Luke, et al.
Published: (2024)
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
Causal Interpretation of Sparse Autoencoder Features in Vision
by: Han, Sangyu, et al.
Published: (2025)
by: Han, Sangyu, et al.
Published: (2025)
A LINDDUN-based Privacy Threat Modeling Framework for GenAI
by: Liao, Qianying, et al.
Published: (2026)
by: Liao, Qianying, et al.
Published: (2026)
Interpretable and Testable Vision Features via Sparse Autoencoders
by: Stevens, Samuel, et al.
Published: (2025)
by: Stevens, Samuel, et al.
Published: (2025)
Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models
by: Swann, Aiden, et al.
Published: (2026)
by: Swann, Aiden, et al.
Published: (2026)
An X-Ray Is Worth 15 Features: Sparse Autoencoders for Interpretable Radiology Report Generation
by: Abdulaal, Ahmed, et al.
Published: (2024)
by: Abdulaal, Ahmed, et al.
Published: (2024)
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
by: Paek, Nathan, et al.
Published: (2025)
by: Paek, Nathan, et al.
Published: (2025)
Chemical composition and some biological activities of marine algae collected in Tunisia
by: F Frikha
Published: (2011)
by: F Frikha
Published: (2011)
Interpretation Gaps in LLM-Assisted Comprehension of Privacy Documents
by: Dewri, Rinku
Published: (2025)
by: Dewri, Rinku
Published: (2025)
The Rogue Scalpel: Activation Steering Compromises LLM Safety
by: Korznikov, Anton, et al.
Published: (2025)
by: Korznikov, Anton, et al.
Published: (2025)
Learnable Faster Kernel-PCA for Nonlinear Fault Detection: Deep Autoencoder-Based Realization
by: Ren, Zelin, et al.
Published: (2021)
by: Ren, Zelin, et al.
Published: (2021)
Water Demand Maximization: Quick Recovery of Nonlinear Physics Solutions
by: Hari, Sai Krishna Kanth, et al.
Published: (2026)
by: Hari, Sai Krishna Kanth, et al.
Published: (2026)
Mechanistic Interpretability of ASR models using Sparse Autoencoders
by: Pluth, Dan, et al.
Published: (2026)
by: Pluth, Dan, et al.
Published: (2026)
Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
by: Du, Hongfei, et al.
Published: (2026)
by: Du, Hongfei, et al.
Published: (2026)
Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders
by: Wu, John F., et al.
Published: (2025)
by: Wu, John F., et al.
Published: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Interpretable Company Similarity with Sparse Autoencoders
by: Molinari, Marco, et al.
Published: (2024)
by: Molinari, Marco, et al.
Published: (2024)
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
A Comparative CFD–PIV Analysis of a Stirred Tank Equipped with a Rushton and a Pitched Blade Turbine
by: Ali Ben Kilani, et al.
Published: (2025)
by: Ali Ben Kilani, et al.
Published: (2025)
Interpreting Differential Privacy in Terms of Disclosure Risk
by: Kazan, Zeki, et al.
Published: (2025)
by: Kazan, Zeki, et al.
Published: (2025)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Privacy-Preserving Autoencoder for Collaborative Object Detection
by: Azizian, Bardia, et al.
Published: (2024)
by: Azizian, Bardia, et al.
Published: (2024)
Rahand Cellular Memory Echo (RCME) Theory – A Structural Hypothesis for Non-Genetic Cellular Memory
by: Ar, Rahand
Published: (2026)
by: Ar, Rahand
Published: (2026)
Robust Skin Color Driven Privacy Preserving Face Recognition via Function Secret Sharing
by: Han, Dong, et al.
Published: (2024)
by: Han, Dong, et al.
Published: (2024)
InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders
by: Simon, Elana, et al.
Published: (2024)
by: Simon, Elana, et al.
Published: (2024)
Interpreting Network Differential Privacy
by: Hehir, Jonathan, et al.
Published: (2025)
by: Hehir, Jonathan, et al.
Published: (2025)
Measuring Sparse Autoencoder Feature Sensitivity
by: Tian, Claire, et al.
Published: (2025)
by: Tian, Claire, et al.
Published: (2025)
Sparse Autoencoder Features for Classifications and Transferability
by: Gallifant, Jack, et al.
Published: (2025)
by: Gallifant, Jack, et al.
Published: (2025)
Interpreting CFD Surrogates through Sparse Autoencoders
by: Hu, Yeping, et al.
Published: (2025)
by: Hu, Yeping, et al.
Published: (2025)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
by: Tolooshams, Bahareh, et al.
Published: (2025)
by: Tolooshams, Bahareh, et al.
Published: (2025)
Interpretable Reward Model via Sparse Autoencoder
by: Zhang, Shuyi, et al.
Published: (2025)
by: Zhang, Shuyi, et al.
Published: (2025)
Similar Items
-
IncogniText: Privacy-enhancing Conditional Text Anonymization via LLM-based Private Attribute Randomization
by: Frikha, Ahmed, et al.
Published: (2024) -
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
by: Nakka, Krishna Kanth, et al.
Published: (2024) -
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
by: Nakka, Krishna Kanth, et al.
Published: (2024) -
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
by: Frikha, Ahmed, et al.
Published: (2024) -
Mammo-SAE: Interpreting Breast Cancer Concept Learning with Sparse Autoencoders
by: Nakka, Krishna Kanth
Published: (2025)