ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Zeming, Wu, Chengcan, Sun, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
by: Zhang, Yihao, et al.
Published: (2026)
by: Zhang, Yihao, et al.
Published: (2026)
Automata-Based Steering of Large Language Models for Diverse Structured Generation
by: Luan, Xiaokun, et al.
Published: (2025)
by: Luan, Xiaokun, et al.
Published: (2025)
MILE: A Mutation Testing Framework of In-Context Learning Systems
by: Wei, Zeming, et al.
Published: (2024)
by: Wei, Zeming, et al.
Published: (2024)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)
by: Kaunismaa, Jackson, et al.
Published: (2026)
SCoPE: Evaluating LLMs for Software Vulnerability Detection
by: Gonçalves, José, et al.
Published: (2024)
by: Gonçalves, José, et al.
Published: (2024)
Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision
by: Mukherjee, Manisha, et al.
Published: (2026)
by: Mukherjee, Manisha, et al.
Published: (2026)
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
by: Asthana, Shubhi, et al.
Published: (2025)
by: Asthana, Shubhi, et al.
Published: (2025)
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
by: Spracklen, Joseph, et al.
Published: (2024)
by: Spracklen, Joseph, et al.
Published: (2024)
Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
by: Zhang, Peng, et al.
Published: (2025)
by: Zhang, Peng, et al.
Published: (2025)
LLMSecConfig: An LLM-Based Approach for Fixing Software Container Misconfigurations
by: Ye, Ziyang, et al.
Published: (2025)
by: Ye, Ziyang, et al.
Published: (2025)
Revisiting Pre-trained Language Models for Vulnerability Detection
by: Li, Youpeng, et al.
Published: (2025)
by: Li, Youpeng, et al.
Published: (2025)
Which PPML Would a User Choose? A Structured Decision Support Framework for Developers to Rank PPML Techniques Based on User Acceptance Criteria
by: Löbner, Sascha, et al.
Published: (2024)
by: Löbner, Sascha, et al.
Published: (2024)
Traversal-as-Policy: Log-Distilled Gated Behavior Trees as Externalized, Verifiable Policies for Safe, Robust, and Efficient Agents
by: Li, Peiran, et al.
Published: (2026)
by: Li, Peiran, et al.
Published: (2026)
Cyber Threat Detection and Vulnerability Assessment System using Generative AI and Large Language Model
by: M, Keerthi Kumar., et al.
Published: (2026)
by: M, Keerthi Kumar., et al.
Published: (2026)
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection
by: Charoenwet, Wachiraphan, et al.
Published: (2026)
by: Charoenwet, Wachiraphan, et al.
Published: (2026)
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
HexaCoder: Secure Code Generation via Oracle-Guided Synthetic Training Data
by: Hajipour, Hossein, et al.
Published: (2024)
by: Hajipour, Hossein, et al.
Published: (2024)
Security smells in infrastructure as code: a taxonomy update beyond the seven sins
by: War, Aicha, et al.
Published: (2025)
by: War, Aicha, et al.
Published: (2025)
Detection of security smells in IaC scripts through semantics-aware code and language processing
by: War, Aicha, et al.
Published: (2025)
by: War, Aicha, et al.
Published: (2025)
A Robust Pipeline for Differentially Private Federated Learning on Imbalanced Clinical Data using SMOTETomek and FedProx
by: Tertulino, Rodrigo
Published: (2025)
by: Tertulino, Rodrigo
Published: (2025)
StepShield: When, Not Whether to Intervene on Rogue Agents
by: Felicia, Gloria, et al.
Published: (2026)
by: Felicia, Gloria, et al.
Published: (2026)
Constrained Decoding for Secure Code Generation
by: Fu, Yanjun, et al.
Published: (2024)
by: Fu, Yanjun, et al.
Published: (2024)
VulCatch: Enhancing Binary Vulnerability Detection through CodeT5 Decompilation and KAN Advanced Feature Extraction
by: Chukkol, Abdulrahman Hamman Adama, et al.
Published: (2024)
by: Chukkol, Abdulrahman Hamman Adama, et al.
Published: (2024)
Defending against Adversarial Malware Attacks on ML-based Android Malware Detection Systems
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
Tady: A Neural Disassembler without Structural Constraint Violations
by: Qin, Siliang, et al.
Published: (2025)
by: Qin, Siliang, et al.
Published: (2025)
Innovations in Cardless Artificial Intelligence Banking: A Comprehensive Framework for Cyber Secure and Fraud Mitigation using Machine Learning Algorithms
by: Israfeel, Md
Published: (2026)
by: Israfeel, Md
Published: (2026)
Evaluating LLaMA 3.2 for Software Vulnerability Detection
by: Gonçalves, José, et al.
Published: (2025)
by: Gonçalves, José, et al.
Published: (2025)
Instruction Tuning for Secure Code Generation
by: He, Jingxuan, et al.
Published: (2024)
by: He, Jingxuan, et al.
Published: (2024)
Prompting Techniques for Secure Code Generation: A Systematic Investigation
by: Tony, Catherine, et al.
Published: (2024)
by: Tony, Catherine, et al.
Published: (2024)
Software Vulnerability Detection Using a Lightweight Graph Neural Network
by: Farmer, Miles, et al.
Published: (2026)
by: Farmer, Miles, et al.
Published: (2026)
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
by: Sun, Yuqiang, et al.
Published: (2024)
by: Sun, Yuqiang, et al.
Published: (2024)
LUNA: A Model-Based Universal Analysis Framework for Large Language Models
by: Song, Da, et al.
Published: (2023)
by: Song, Da, et al.
Published: (2023)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
by: Hubinger, Evan, et al.
Published: (2024)
by: Hubinger, Evan, et al.
Published: (2024)
Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward
by: Xie, Xuan, et al.
Published: (2024)
by: Xie, Xuan, et al.
Published: (2024)
Rethinking and Exploring String-Based Malware Family Classification in the Era of LLMs and RAG
by: Chen, Yufan, et al.
Published: (2025)
by: Chen, Yufan, et al.
Published: (2025)
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
by: Ren, Qibing, et al.
Published: (2024)
by: Ren, Qibing, et al.
Published: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
SPDZCoder: Combining Expert Knowledge with LLMs for Generating Privacy-Computing Code
by: Dong, Xiaoning, et al.
Published: (2024)
by: Dong, Xiaoning, et al.
Published: (2024)
Similar Items
-
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
by: Wei, Zeming, et al.
Published: (2026) -
ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
by: Zhang, Yihao, et al.
Published: (2026) -
Automata-Based Steering of Large Language Models for Diverse Structured Generation
by: Luan, Xiaokun, et al.
Published: (2025) -
MILE: A Mutation Testing Framework of In-Context Learning Systems
by: Wei, Zeming, et al.
Published: (2024) -
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)