PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
Fuente:
arXiv
Saved in:
| Main Authors: | Manczak, Blazej, Zemour, Eliott, Lin, Eric, Mugunthan, Vaikkunth |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
by: Eiras, Francisco, et al.
Published: (2025)
by: Eiras, Francisco, et al.
Published: (2025)
AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning
by: Gu, Shanzhi, et al.
Published: (2025)
by: Gu, Shanzhi, et al.
Published: (2025)
Retrieval-Augmented Few-Shot Prompting Versus Fine-Tuning for Code Vulnerability Detection
by: Trad, Fouad, et al.
Published: (2025)
by: Trad, Fouad, et al.
Published: (2025)
DeepGuard: Secure Code Generation via Multi-Layer Semantic Aggregation
by: Huang, Li, et al.
Published: (2026)
by: Huang, Li, et al.
Published: (2026)
A Systematic Approach to Predict the Impact of Cybersecurity Vulnerabilities Using LLMs
by: Høst, Anders Mølmen, et al.
Published: (2025)
by: Høst, Anders Mølmen, et al.
Published: (2025)
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
R1-Fuzz: Specializing Language Models for Textual Fuzzing via Reinforcement Learning
by: Lin, Jiayi, et al.
Published: (2025)
by: Lin, Jiayi, et al.
Published: (2025)
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
by: Wang, Su, et al.
Published: (2026)
by: Wang, Su, et al.
Published: (2026)
LLMs + Security = Trouble
by: Livshits, Benjamin
Published: (2026)
by: Livshits, Benjamin
Published: (2026)
LLMs as verification oracles for Solidity
by: Bartoletti, Massimo, et al.
Published: (2025)
by: Bartoletti, Massimo, et al.
Published: (2025)
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)
by: Kaunismaa, Jackson, et al.
Published: (2026)
Harnessing the Power of LLMs in Source Code Vulnerability Detection
by: Mahyari, Andrew A
Published: (2024)
by: Mahyari, Andrew A
Published: (2024)
Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities
by: Garg, Aayush, et al.
Published: (2025)
by: Garg, Aayush, et al.
Published: (2025)
SPDZCoder: Combining Expert Knowledge with LLMs for Generating Privacy-Computing Code
by: Dong, Xiaoning, et al.
Published: (2024)
by: Dong, Xiaoning, et al.
Published: (2024)
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection
by: Zhao, Zijie, et al.
Published: (2026)
by: Zhao, Zijie, et al.
Published: (2026)
Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks
by: Qu, Yubin, et al.
Published: (2026)
by: Qu, Yubin, et al.
Published: (2026)
Arbiter: Detecting Interference in LLM Agent System Prompts
by: Mason, Tony
Published: (2026)
by: Mason, Tony
Published: (2026)
Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic Transformation
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025)
by: Chen, Junkai, et al.
Published: (2025)
Evaluating and Mitigating Linguistic Discrimination in Large Language Models
by: Dong, Guoliang, et al.
Published: (2024)
by: Dong, Guoliang, et al.
Published: (2024)
Web Agents Should Adopt the Plan-Then-Execute Paradigm
by: Piet, Julien, et al.
Published: (2026)
by: Piet, Julien, et al.
Published: (2026)
ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse
by: Liu, Simiao, et al.
Published: (2026)
by: Liu, Simiao, et al.
Published: (2026)
Can LLM Prompting Serve as a Proxy for Static Analysis in Vulnerability Detection
by: Ceka, Ira, et al.
Published: (2024)
by: Ceka, Ira, et al.
Published: (2024)
Exploring ChatGPT's Capabilities on Vulnerability Management
by: Liu, Peiyu, et al.
Published: (2023)
by: Liu, Peiyu, et al.
Published: (2023)
Fuzzing: Randomness? Reasoning! Efficient Directed Fuzzing via Large Language Models
by: Feng, Xiaotao, et al.
Published: (2025)
by: Feng, Xiaotao, et al.
Published: (2025)
Efficient Detection of Toxic Prompts in Large Language Models
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Prompt Injection attack against LLM-integrated Applications
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
Agentic Specification Generator for Move Programs
by: Fu, Yu-Fu, et al.
Published: (2025)
by: Fu, Yu-Fu, et al.
Published: (2025)
Beyond Classification: Evaluating LLMs for Fine-Grained Automatic Malware Behavior Auditing
by: Zheng, Xinran, et al.
Published: (2025)
by: Zheng, Xinran, et al.
Published: (2025)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
by: Shen, Qingchao, et al.
Published: (2026)
by: Shen, Qingchao, et al.
Published: (2026)
Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
by: Yuan, He Yang, et al.
Published: (2026)
by: Yuan, He Yang, et al.
Published: (2026)
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
by: Dubniczky, Richard A., et al.
Published: (2025)
by: Dubniczky, Richard A., et al.
Published: (2025)
Rethinking and Exploring String-Based Malware Family Classification in the Era of LLMs and RAG
by: Chen, Yufan, et al.
Published: (2025)
by: Chen, Yufan, et al.
Published: (2025)
How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code?
by: Yang, Hua, et al.
Published: (2025)
by: Yang, Hua, et al.
Published: (2025)
Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented Scanning
by: Yan, Shenao, et al.
Published: (2026)
by: Yan, Shenao, et al.
Published: (2026)
Similar Items
-
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
by: Eiras, Francisco, et al.
Published: (2025) -
AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
by: Yang, Rui, et al.
Published: (2025) -
Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning
by: Gu, Shanzhi, et al.
Published: (2025) -
Retrieval-Augmented Few-Shot Prompting Versus Fine-Tuning for Code Vulnerability Detection
by: Trad, Fouad, et al.
Published: (2025) -
DeepGuard: Secure Code Generation via Multi-Layer Semantic Aggregation
by: Huang, Li, et al.
Published: (2026)