SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Piras, Giorgio, Mura, Raffaele, Brau, Fabio, Oneto, Luca, Roli, Fabio, Biggio, Battista |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent-space Attacks for Refusal Evasion in Language Models
by: Piras, Giorgio, et al.
Published: (2026)
by: Piras, Giorgio, et al.
Published: (2026)
LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
by: Mura, Raffaele, et al.
Published: (2025)
by: Mura, Raffaele, et al.
Published: (2025)
S2AP: Score-space Sharpness Minimization for Adversarial Pruning
by: Piras, Giorgio, et al.
Published: (2025)
by: Piras, Giorgio, et al.
Published: (2025)
Exploiting Edge Features for Transferable Adversarial Attacks in Distributed Machine Learning
by: Rossolini, Giulio, et al.
Published: (2025)
by: Rossolini, Giulio, et al.
Published: (2025)
HO-FMN: Hyperparameter Optimization for Fast Minimum-Norm Attacks
by: Mura, Raffaele, et al.
Published: (2024)
by: Mura, Raffaele, et al.
Published: (2024)
SAGE-5GC: Security-Aware Guidelines for Evaluating Anomaly Detection in the 5G Core Network
by: Manca, Cristian, et al.
Published: (2026)
by: Manca, Cristian, et al.
Published: (2026)
Adversarial Pruning: A Survey and Benchmark of Pruning Methods for Adversarial Robustness
by: Piras, Giorgio, et al.
Published: (2024)
by: Piras, Giorgio, et al.
Published: (2024)
Demystifying the Role of Rule-based Detection in AI Systems for Windows Malware Detection
by: Ponte, Andrea, et al.
Published: (2025)
by: Ponte, Andrea, et al.
Published: (2025)
Robustness-Congruent Adversarial Training for Secure Machine Learning Model Updates
by: Angioni, Daniele, et al.
Published: (2024)
by: Angioni, Daniele, et al.
Published: (2024)
Evaluating the Evaluators: Trust in Adversarial Robustness Tests
by: Cinà, Antonio Emanuele, et al.
Published: (2025)
by: Cinà, Antonio Emanuele, et al.
Published: (2025)
Nebula: Self-Attention for Dynamic Malware Analysis
by: Trizna, Dmitrijs, et al.
Published: (2023)
by: Trizna, Dmitrijs, et al.
Published: (2023)
Robust Synthetic Data-Driven Detection of Living-Off-the-Land Reverse Shells
by: Trizna, Dmitrijs, et al.
Published: (2024)
by: Trizna, Dmitrijs, et al.
Published: (2024)
Buffer-free Class-Incremental Learning with Out-of-Distribution Detection
by: Gupta, Srishti, et al.
Published: (2025)
by: Gupta, Srishti, et al.
Published: (2025)
On the Robustness of Adversarial Training Against Uncertainty Attacks
by: Ledda, Emanuele, et al.
Published: (2024)
by: Ledda, Emanuele, et al.
Published: (2024)
Refusal in Language Models Is Mediated by a Single Direction
by: Arditi, Andy, et al.
Published: (2024)
by: Arditi, Andy, et al.
Published: (2024)
Regression-aware Continual Learning for Android Malware Detection
by: Ghiani, Daniele, et al.
Published: (2025)
by: Ghiani, Daniele, et al.
Published: (2025)
On the Failure of Topic-Matched Contrast Baselines in Multi-Directional Refusal Abliteration
by: Petrov, Valentin
Published: (2026)
by: Petrov, Valentin
Published: (2026)
A Neural Rejection System Against Universal Adversarial Perturbations in Radio Signal Classification
by: Zhang, Lu, et al.
Published: (2025)
by: Zhang, Lu, et al.
Published: (2025)
SLIFER: Investigating Performance and Robustness of Malware Detection Pipelines
by: Ponte, Andrea, et al.
Published: (2024)
by: Ponte, Andrea, et al.
Published: (2024)
Evaluating Line-level Localization Ability of Learning-based Code Vulnerability Detection Models
by: Pintore, Marco, et al.
Published: (2025)
by: Pintore, Marco, et al.
Published: (2025)
Out-of-Distribution Detection for Continual Learning: Design Principles and Benchmarking
by: Gupta, Srishti, et al.
Published: (2025)
by: Gupta, Srishti, et al.
Published: (2025)
Two Heads Are Better than One: Simulating Large Transformers with Small Ones
by: Yu, Hantao, et al.
Published: (2025)
by: Yu, Hantao, et al.
Published: (2025)
Energy-Latency Attacks via Sponge Poisoning
by: Cinà, Antonio Emanuele, et al.
Published: (2022)
by: Cinà, Antonio Emanuele, et al.
Published: (2022)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Empirical Quantification of Spurious Correlations in Malware Detection
by: Perasso, Bianca, et al.
Published: (2025)
by: Perasso, Bianca, et al.
Published: (2025)
Towards One-shot Federated Learning: Advances, Challenges, and Future Directions
by: Amato, Flora, et al.
Published: (2025)
by: Amato, Flora, et al.
Published: (2025)
Attention-based Adversarial Robust Distillation in Radio Signal Classifications for Low-Power IoT Devices
by: Zhang, Lu, et al.
Published: (2025)
by: Zhang, Lu, et al.
Published: (2025)
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
by: Rahimi, Eliron, et al.
Published: (2026)
by: Rahimi, Eliron, et al.
Published: (2026)
Hyperbolic Learning with Multimodal Large Language Models
by: Mandica, Paolo, et al.
Published: (2024)
by: Mandica, Paolo, et al.
Published: (2024)
Over-parameterization and Adversarial Robustness in Neural Networks: An Overview and Empirical Analysis
by: Gupta, Srishti, et al.
Published: (2024)
by: Gupta, Srishti, et al.
Published: (2024)
ImageNet-Patch: A Dataset for Benchmarking Machine Learning Robustness against Adversarial Patches
by: Pintor, Maura, et al.
Published: (2022)
by: Pintor, Maura, et al.
Published: (2022)
On the Bias of Next-Token Predictors Toward Systematically Inefficient Reasoning: A Shortest-Path Case Study
by: Alberghi, Riccardo, et al.
Published: (2025)
by: Alberghi, Riccardo, et al.
Published: (2025)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
DirectMultiStep: Direct Route Generation for Multistep Retrosynthesis
by: Shee, Yu, et al.
Published: (2024)
by: Shee, Yu, et al.
Published: (2024)
Model-Based Reinforcement Learning in Discrete-Action Non-Markovian Reward Decision Processes
by: Trapasso, Alessandro, et al.
Published: (2025)
by: Trapasso, Alessandro, et al.
Published: (2025)
CIVeX: Causal Intervention Verification for Language Agents
by: Rovai, Fabio
Published: (2026)
by: Rovai, Fabio
Published: (2026)
Two Heads Are Better than One: Model-Weight and Latent-Space Analysis for Federated Learning on Non-iid Data against Poisoning Attacks
by: Lyu, Xingyu, et al.
Published: (2025)
by: Lyu, Xingyu, et al.
Published: (2025)
Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning
by: Yang, Puning, et al.
Published: (2026)
by: Yang, Puning, et al.
Published: (2026)
From Coefficients to Directions: Rethinking Model Merging with Directional Alignment
by: Chen, Zhikang, et al.
Published: (2025)
by: Chen, Zhikang, et al.
Published: (2025)
SatSOM: Saturation Self-Organizing Maps for Continual Learning
by: Urbanik, Igor, et al.
Published: (2025)
by: Urbanik, Igor, et al.
Published: (2025)
Similar Items
-
Latent-space Attacks for Refusal Evasion in Language Models
by: Piras, Giorgio, et al.
Published: (2026) -
LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
by: Mura, Raffaele, et al.
Published: (2025) -
S2AP: Score-space Sharpness Minimization for Adversarial Pruning
by: Piras, Giorgio, et al.
Published: (2025) -
Exploiting Edge Features for Transferable Adversarial Attacks in Distributed Machine Learning
by: Rossolini, Giulio, et al.
Published: (2025) -
HO-FMN: Hyperparameter Optimization for Fast Minimum-Norm Attacks
by: Mura, Raffaele, et al.
Published: (2024)