LLMGuard: Guarding Against Unsafe LLM Behavior
Fuente:
arXiv
Guardado en:
| Autores principales: | Goyal, Shubh, Hira, Medha, Mishra, Shubham, Goyal, Sukriti, Goel, Arnav, Dadu, Niharika, DB, Kirushikesh, Mehta, Sameep, Madaan, Nishtha |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unclonable Cryptography with Unbounded Collusions and Impossibility of Hyperefficient Shadow Tomography
por: Çakan, Alper, et al.
Publicado: (2023)
por: Çakan, Alper, et al.
Publicado: (2023)
Proofs of No Intrusion
por: Goyal, Vipul, et al.
Publicado: (2025)
por: Goyal, Vipul, et al.
Publicado: (2025)
How to Copy-Protect Malleable-Puncturable Cryptographic Functionalities Under Arbitrary Challenge Distributions
por: Çakan, Alper, et al.
Publicado: (2025)
por: Çakan, Alper, et al.
Publicado: (2025)
CircuitGuard: Mitigating LLM Memorization in RTL Code Generation Against IP Leakage
por: Mashnoor, Nowfel, et al.
Publicado: (2025)
por: Mashnoor, Nowfel, et al.
Publicado: (2025)
MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
por: Wang, Zhiqiang, et al.
Publicado: (2025)
por: Wang, Zhiqiang, et al.
Publicado: (2025)
Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts
por: Hasan, Md. Mehedi, et al.
Publicado: (2025)
por: Hasan, Md. Mehedi, et al.
Publicado: (2025)
GuardPhish: Securing Open-Source LLMs from Phishing Abuse
por: Mishra, Rina, et al.
Publicado: (2026)
por: Mishra, Rina, et al.
Publicado: (2026)
DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation
por: Jiang, Bo
Publicado: (2026)
por: Jiang, Bo
Publicado: (2026)
ORCHID: Streaming Threat Detection over Versioned Provenance Graphs
por: Goyal, Akul, et al.
Publicado: (2024)
por: Goyal, Akul, et al.
Publicado: (2024)
SAGE: Sample-Aware Guarding Engine for Robust Intrusion Detection Against Adversarial Attacks
por: Chen, Jing, et al.
Publicado: (2025)
por: Chen, Jing, et al.
Publicado: (2025)
RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing
por: Zhang, Wenhui, et al.
Publicado: (2026)
por: Zhang, Wenhui, et al.
Publicado: (2026)
An Undeniable Signature Scheme Utilizing Module Lattices
por: Dey, Kunal, et al.
Publicado: (2024)
por: Dey, Kunal, et al.
Publicado: (2024)
Propagating Unsafe Actions in LLM Controlled Multi-Robot Collaboration via Single Robot Compromise
por: Huang, Zhen, et al.
Publicado: (2026)
por: Huang, Zhen, et al.
Publicado: (2026)
KinGuard: Hierarchical Kinship-Aware Fingerprinting to Defend Against Large Language Model Stealing
por: Xu, Zhenhua, et al.
Publicado: (2026)
por: Xu, Zhenhua, et al.
Publicado: (2026)
LLM Security Guard for Code
por: Kavian, Arya, et al.
Publicado: (2024)
por: Kavian, Arya, et al.
Publicado: (2024)
Targeted Fuzzing for Unsafe Rust Code: Leveraging Selective Instrumentation
por: Paaßen, David, et al.
Publicado: (2025)
por: Paaßen, David, et al.
Publicado: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
por: Yuan, Lingzhi, et al.
Publicado: (2025)
por: Yuan, Lingzhi, et al.
Publicado: (2025)
GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
por: Xiang, Yuxiao, et al.
Publicado: (2025)
por: Xiang, Yuxiao, et al.
Publicado: (2025)
Semantic Differentiation for Tackling Challenges in Watermarking Low-Entropy Constrained Generation Outputs
por: Le, Nghia T., et al.
Publicado: (2026)
por: Le, Nghia T., et al.
Publicado: (2026)
How to Delete Without a Trace: Certified Deniability in a Quantum World
por: Çakan, Alper, et al.
Publicado: (2024)
por: Çakan, Alper, et al.
Publicado: (2024)
Public-Key Quantum Fire and Key-Fire From Classical Oracles
por: Çakan, Alper, et al.
Publicado: (2025)
por: Çakan, Alper, et al.
Publicado: (2025)
Anonymous Public-Key Quantum Money and Quantum Voting
por: Cakan, Alper, et al.
Publicado: (2024)
por: Cakan, Alper, et al.
Publicado: (2024)
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
por: Zhao, Wei, et al.
Publicado: (2026)
por: Zhao, Wei, et al.
Publicado: (2026)
NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks
por: Asl, Javad Rafiei, et al.
Publicado: (2025)
por: Asl, Javad Rafiei, et al.
Publicado: (2025)
MirGuard: Towards a Robust Provenance-based Intrusion Detection System Against Graph Manipulation Attacks
por: Sang, Anyuan, et al.
Publicado: (2025)
por: Sang, Anyuan, et al.
Publicado: (2025)
GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy
por: Minko, Bogdan, et al.
Publicado: (2026)
por: Minko, Bogdan, et al.
Publicado: (2026)
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
por: Zhang, Xiaoyu, et al.
Publicado: (2023)
por: Zhang, Xiaoyu, et al.
Publicado: (2023)
Enabling High-Frequency Trading with Near-Instant, Trustless Cross-Chain Transactions via Pre-Signing Adaptor Signatures
por: Francolla, Ethan, et al.
Publicado: (2025)
por: Francolla, Ethan, et al.
Publicado: (2025)
Securing Genomic Data Against Inference Attacks in Federated Learning Environments
por: Pathade, Chetan, et al.
Publicado: (2025)
por: Pathade, Chetan, et al.
Publicado: (2025)
From Stealthy Data Fabrication to Unsafe Driving: Realistic Scenario Attacks on Collaborative Perception
por: Zhang, Qingzhao, et al.
Publicado: (2026)
por: Zhang, Qingzhao, et al.
Publicado: (2026)
Fast Summary-based Whole-program Analysis to Identify Unsafe Memory Accesses in Rust
por: Zhou, Jie, et al.
Publicado: (2023)
por: Zhou, Jie, et al.
Publicado: (2023)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
por: Wu, Yixin, et al.
Publicado: (2023)
por: Wu, Yixin, et al.
Publicado: (2023)
Hacking, The Lazy Way: LLM Augmented Pentesting
por: Goyal, Dhruva, et al.
Publicado: (2024)
por: Goyal, Dhruva, et al.
Publicado: (2024)
AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent
por: Luo, Jiaqi, et al.
Publicado: (2026)
por: Luo, Jiaqi, et al.
Publicado: (2026)
SpaLLM-Guard: Pairing SMS Spam Detection Using Open-source and Commercial LLMs
por: Salman, Muhammad, et al.
Publicado: (2025)
por: Salman, Muhammad, et al.
Publicado: (2025)
Quantum Machine Learning for Cybersecurity: A Taxonomy and Future Directions
por: Sai, Siva, et al.
Publicado: (2025)
por: Sai, Siva, et al.
Publicado: (2025)
Unclonable Secret Sharing
por: Ananth, Prabhanjan, et al.
Publicado: (2024)
por: Ananth, Prabhanjan, et al.
Publicado: (2024)
Quantum Key Leasing for PKE and FHE with a Classical Lessor
por: Chardouvelis, Orestis, et al.
Publicado: (2023)
por: Chardouvelis, Orestis, et al.
Publicado: (2023)
Delay-Induced Watermarking for Detection of Replay Attacks in Linear Systems
por: Somarakis, Christoforos, et al.
Publicado: (2024)
por: Somarakis, Christoforos, et al.
Publicado: (2024)
GLiGuard: Schema-Conditioned Classification for LLM Safeguard
por: Zaratiana, Urchade, et al.
Publicado: (2026)
por: Zaratiana, Urchade, et al.
Publicado: (2026)
Ejemplares similares
-
Unclonable Cryptography with Unbounded Collusions and Impossibility of Hyperefficient Shadow Tomography
por: Çakan, Alper, et al.
Publicado: (2023) -
Proofs of No Intrusion
por: Goyal, Vipul, et al.
Publicado: (2025) -
How to Copy-Protect Malleable-Puncturable Cryptographic Functionalities Under Arbitrary Challenge Distributions
por: Çakan, Alper, et al.
Publicado: (2025) -
CircuitGuard: Mitigating LLM Memorization in RTL Code Generation Against IP Leakage
por: Mashnoor, Nowfel, et al.
Publicado: (2025) -
MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
por: Wang, Zhiqiang, et al.
Publicado: (2025)