Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
Fuente:
arXiv
Salvato in:
| Autori principali: | Das, Nilanjana, Gaur, Manas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
Human-Interpretable Adversarial Prompt Attack on Large Language Models with Situational Context
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
di: Zhang, Ruikang, et al.
Pubblicazione: (2026)
di: Zhang, Ruikang, et al.
Pubblicazione: (2026)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
di: Haider, Batool, et al.
Pubblicazione: (2025)
di: Haider, Batool, et al.
Pubblicazione: (2025)
Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2026)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2026)
SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
Exploring the Personality Traits of LLMs through Latent Features Steering
di: Yang, Shu, et al.
Pubblicazione: (2024)
di: Yang, Shu, et al.
Pubblicazione: (2024)
Neurosymbolic Retrievers for Retrieval-augmented Generation
di: Saxena, Yash, et al.
Pubblicazione: (2026)
di: Saxena, Yash, et al.
Pubblicazione: (2026)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
di: Lamb, Tom A., et al.
Pubblicazione: (2024)
di: Lamb, Tom A., et al.
Pubblicazione: (2024)
Exploitation Without Deception: Dark Triad Feature Steering Reveals Separable Antisocial Circuits in Language Models
di: Berg, Cameron, et al.
Pubblicazione: (2026)
di: Berg, Cameron, et al.
Pubblicazione: (2026)
Towards Inference-time Category-wise Safety Steering for Large Language Models
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
di: Saha, Anusa, et al.
Pubblicazione: (2026)
di: Saha, Anusa, et al.
Pubblicazione: (2026)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
di: Shelat, Shlok, et al.
Pubblicazione: (2026)
di: Shelat, Shlok, et al.
Pubblicazione: (2026)
From Guessing to Asking: An Approach to Resolving the Persona Knowledge Gap in LLMs during Multi-Turn Conversations
di: Baskar, Sarvesh, et al.
Pubblicazione: (2025)
di: Baskar, Sarvesh, et al.
Pubblicazione: (2025)
Layer-wise Regularized Dropout for Neural Language Models
di: Ni, Shiwen, et al.
Pubblicazione: (2024)
di: Ni, Shiwen, et al.
Pubblicazione: (2024)
SaGE: Evaluating Moral Consistency in Large Language Models
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024)
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
di: Joshi, Abhinav, et al.
Pubblicazione: (2024)
di: Joshi, Abhinav, et al.
Pubblicazione: (2024)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2025)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2025)
CogSteer: Cognition-Inspired Selective Layer Intervention for Efficiently Steering Large Language Models
di: Wang, Xintong, et al.
Pubblicazione: (2024)
di: Wang, Xintong, et al.
Pubblicazione: (2024)
DESTEIN: Navigating Detoxification of Language Models via Universal Steering Pairs and Head-wise Activation Fusion
di: Li, Yu, et al.
Pubblicazione: (2024)
di: Li, Yu, et al.
Pubblicazione: (2024)
Layer-wise Positional Bias in Short-Context Language Modeling
di: Rahimi, Maryam, et al.
Pubblicazione: (2026)
di: Rahimi, Maryam, et al.
Pubblicazione: (2026)
Unsupervised Layer-wise Score Aggregation for Textual OOD Detection
di: Darrin, Maxime, et al.
Pubblicazione: (2023)
di: Darrin, Maxime, et al.
Pubblicazione: (2023)
IoT-Based Preventive Mental Health Using Knowledge Graphs and Standards for Better Well-Being
di: Gyrard, Amelie, et al.
Pubblicazione: (2024)
di: Gyrard, Amelie, et al.
Pubblicazione: (2024)
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2024)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2024)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)
KV Cache Steering for Controlling Frozen LLMs
di: Belitsky, Max, et al.
Pubblicazione: (2025)
di: Belitsky, Max, et al.
Pubblicazione: (2025)
Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
di: Walkowiak, Paweł, et al.
Pubblicazione: (2025)
di: Walkowiak, Paweł, et al.
Pubblicazione: (2025)
Evolutionary Feature-wise Thresholding for Binary Representation of NLP Embeddings
di: Sinha, Soumen, et al.
Pubblicazione: (2025)
di: Sinha, Soumen, et al.
Pubblicazione: (2025)
Efficient Layer-wise LLM Fine-tuning for Revision Intention Prediction
di: Liu, Zhexiong, et al.
Pubblicazione: (2025)
di: Liu, Zhexiong, et al.
Pubblicazione: (2025)
Steering Towards Fairness: Mitigating Political Bias in LLMs
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
di: Pres, Itamar, et al.
Pubblicazione: (2024)
di: Pres, Itamar, et al.
Pubblicazione: (2024)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
di: Chandna, Bhavik, et al.
Pubblicazione: (2025)
di: Chandna, Bhavik, et al.
Pubblicazione: (2025)
Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
di: Lian, Jiawei, et al.
Pubblicazione: (2025)
di: Lian, Jiawei, et al.
Pubblicazione: (2025)
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
di: Mohseni, Seyedreza, et al.
Pubblicazione: (2024)
di: Mohseni, Seyedreza, et al.
Pubblicazione: (2024)
A Framework to Assess Multilingual Vulnerabilities of LLMs
di: Tang, Likai, et al.
Pubblicazione: (2025)
di: Tang, Likai, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of GPT-2: Lexical and Contextual Layers in Sentiment Analysis
di: Hatua, Amartya
Pubblicazione: (2025)
di: Hatua, Amartya
Pubblicazione: (2025)
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
di: Zhou, Hanhan, et al.
Pubblicazione: (2026)
di: Zhou, Hanhan, et al.
Pubblicazione: (2026)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
di: Raimondi, Bianca, et al.
Pubblicazione: (2025)
di: Raimondi, Bianca, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
di: Das, Nilanjana, et al.
Pubblicazione: (2024) -
Human-Interpretable Adversarial Prompt Attack on Large Language Models with Situational Context
di: Das, Nilanjana, et al.
Pubblicazione: (2024) -
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
di: Zhang, Ruikang, et al.
Pubblicazione: (2026) -
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
di: Haider, Batool, et al.
Pubblicazione: (2025) -
Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2026)