Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Peng, Sun, Peijie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision
by: Mukherjee, Manisha, et al.
Published: (2026)
by: Mukherjee, Manisha, et al.
Published: (2026)
A Robust Pipeline for Differentially Private Federated Learning on Imbalanced Clinical Data using SMOTETomek and FedProx
by: Tertulino, Rodrigo
Published: (2025)
by: Tertulino, Rodrigo
Published: (2025)
LLMSecConfig: An LLM-Based Approach for Fixing Software Container Misconfigurations
by: Ye, Ziyang, et al.
Published: (2025)
by: Ye, Ziyang, et al.
Published: (2025)
Innovations in Cardless Artificial Intelligence Banking: A Comprehensive Framework for Cyber Secure and Fraud Mitigation using Machine Learning Algorithms
by: Israfeel, Md
Published: (2026)
by: Israfeel, Md
Published: (2026)
MILE: A Mutation Testing Framework of In-Context Learning Systems
by: Wei, Zeming, et al.
Published: (2024)
by: Wei, Zeming, et al.
Published: (2024)
Which PPML Would a User Choose? A Structured Decision Support Framework for Developers to Rank PPML Techniques Based on User Acceptance Criteria
by: Löbner, Sascha, et al.
Published: (2024)
by: Löbner, Sascha, et al.
Published: (2024)
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction
by: Wei, Zeming, et al.
Published: (2025)
by: Wei, Zeming, et al.
Published: (2025)
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Containment Verification: AI Safety Guarantees Independent of Alignment
by: Moon, Royce, et al.
Published: (2026)
by: Moon, Royce, et al.
Published: (2026)
Tady: A Neural Disassembler without Structural Constraint Violations
by: Qin, Siliang, et al.
Published: (2025)
by: Qin, Siliang, et al.
Published: (2025)
ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
by: Zhang, Yihao, et al.
Published: (2026)
by: Zhang, Yihao, et al.
Published: (2026)
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
by: Asthana, Shubhi, et al.
Published: (2025)
by: Asthana, Shubhi, et al.
Published: (2025)
Traversal-as-Policy: Log-Distilled Gated Behavior Trees as Externalized, Verifiable Policies for Safe, Robust, and Efficient Agents
by: Li, Peiran, et al.
Published: (2026)
by: Li, Peiran, et al.
Published: (2026)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
Prompting Techniques for Secure Code Generation: A Systematic Investigation
by: Tony, Catherine, et al.
Published: (2024)
by: Tony, Catherine, et al.
Published: (2024)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
by: Sun, Yuqiang, et al.
Published: (2024)
by: Sun, Yuqiang, et al.
Published: (2024)
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
by: Spracklen, Joseph, et al.
Published: (2024)
by: Spracklen, Joseph, et al.
Published: (2024)
AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Security smells in infrastructure as code: a taxonomy update beyond the seven sins
by: War, Aicha, et al.
Published: (2025)
by: War, Aicha, et al.
Published: (2025)
Detection of security smells in IaC scripts through semantics-aware code and language processing
by: War, Aicha, et al.
Published: (2025)
by: War, Aicha, et al.
Published: (2025)
Defending against Adversarial Malware Attacks on ML-based Android Malware Detection Systems
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
Evaluating LLaMA 3.2 for Software Vulnerability Detection
by: Gonçalves, José, et al.
Published: (2025)
by: Gonçalves, José, et al.
Published: (2025)
Revisiting Pre-trained Language Models for Vulnerability Detection
by: Li, Youpeng, et al.
Published: (2025)
by: Li, Youpeng, et al.
Published: (2025)
StepShield: When, Not Whether to Intervene on Rogue Agents
by: Felicia, Gloria, et al.
Published: (2026)
by: Felicia, Gloria, et al.
Published: (2026)
Constrained Decoding for Secure Code Generation
by: Fu, Yanjun, et al.
Published: (2024)
by: Fu, Yanjun, et al.
Published: (2024)
VulCatch: Enhancing Binary Vulnerability Detection through CodeT5 Decompilation and KAN Advanced Feature Extraction
by: Chukkol, Abdulrahman Hamman Adama, et al.
Published: (2024)
by: Chukkol, Abdulrahman Hamman Adama, et al.
Published: (2024)
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection
by: Charoenwet, Wachiraphan, et al.
Published: (2026)
by: Charoenwet, Wachiraphan, et al.
Published: (2026)
Cyber Threat Detection and Vulnerability Assessment System using Generative AI and Large Language Model
by: M, Keerthi Kumar., et al.
Published: (2026)
by: M, Keerthi Kumar., et al.
Published: (2026)
Instruction Tuning for Secure Code Generation
by: He, Jingxuan, et al.
Published: (2024)
by: He, Jingxuan, et al.
Published: (2024)
SCoPE: Evaluating LLMs for Software Vulnerability Detection
by: Gonçalves, José, et al.
Published: (2024)
by: Gonçalves, José, et al.
Published: (2024)
Software Vulnerability Detection Using a Lightweight Graph Neural Network
by: Farmer, Miles, et al.
Published: (2026)
by: Farmer, Miles, et al.
Published: (2026)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
by: Hubinger, Evan, et al.
Published: (2024)
by: Hubinger, Evan, et al.
Published: (2024)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)
by: Peng, Jiaren, et al.
Published: (2026)
Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward
by: Xie, Xuan, et al.
Published: (2024)
by: Xie, Xuan, et al.
Published: (2024)
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
by: Ren, Qibing, et al.
Published: (2024)
by: Ren, Qibing, et al.
Published: (2024)
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
CASCADE: LLM-Powered JavaScript Deobfuscator at Google
by: Jiang, Shan, et al.
Published: (2025)
by: Jiang, Shan, et al.
Published: (2025)
LUNA: A Model-Based Universal Analysis Framework for Large Language Models
by: Song, Da, et al.
Published: (2023)
by: Song, Da, et al.
Published: (2023)
Similar Items
-
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
by: Wei, Zeming, et al.
Published: (2026) -
Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision
by: Mukherjee, Manisha, et al.
Published: (2026) -
A Robust Pipeline for Differentially Private Federated Learning on Imbalanced Clinical Data using SMOTETomek and FedProx
by: Tertulino, Rodrigo
Published: (2025) -
LLMSecConfig: An LLM-Based Approach for Fixing Software Container Misconfigurations
by: Ye, Ziyang, et al.
Published: (2025) -
Innovations in Cardless Artificial Intelligence Banking: A Comprehensive Framework for Cyber Secure and Fraud Mitigation using Machine Learning Algorithms
by: Israfeel, Md
Published: (2026)