SAGE: A Generic Framework for LLM Safety Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jindal, Madhur, Shrawgi, Hari, Agrawal, Parag, Dandapat, Sandipan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM Safety for Children
von: Rath, Prasanjit, et al.
Veröffentlicht: (2025)
von: Rath, Prasanjit, et al.
Veröffentlicht: (2025)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024)
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024)
LLM-Safety Evaluations Lack Robustness
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
Security of and by Generative AI platforms
von: Hayagreevan, Hari, et al.
Veröffentlicht: (2024)
von: Hayagreevan, Hari, et al.
Veröffentlicht: (2024)
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
von: Ji, Zimo, et al.
Veröffentlicht: (2025)
von: Ji, Zimo, et al.
Veröffentlicht: (2025)
aiXamine: Simplified LLM Safety and Security
von: Deniz, Fatih, et al.
Veröffentlicht: (2025)
von: Deniz, Fatih, et al.
Veröffentlicht: (2025)
FedSecureFormer: A Fast, Federated and Secure Transformer Framework for Lightweight Intrusion Detection in Connected and Autonomous Vehicles
von: S, Devika, et al.
Veröffentlicht: (2025)
von: S, Devika, et al.
Veröffentlicht: (2025)
Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks
von: Sahu, Anubhab, et al.
Veröffentlicht: (2026)
von: Sahu, Anubhab, et al.
Veröffentlicht: (2026)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
von: Kumar, Shanu, et al.
Veröffentlicht: (2024)
von: Kumar, Shanu, et al.
Veröffentlicht: (2024)
FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment
von: Kuznetsov, Daniel, et al.
Veröffentlicht: (2026)
von: Kuznetsov, Daniel, et al.
Veröffentlicht: (2026)
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
von: Li, Xinghang, et al.
Veröffentlicht: (2025)
von: Li, Xinghang, et al.
Veröffentlicht: (2025)
Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks
von: De Stefano, Gianluca, et al.
Veröffentlicht: (2024)
von: De Stefano, Gianluca, et al.
Veröffentlicht: (2024)
Cisco Integrated AI Security and Safety Framework Report
von: Chang, Amy, et al.
Veröffentlicht: (2025)
von: Chang, Amy, et al.
Veröffentlicht: (2025)
Safety Layers in Aligned Large Language Models: The Key to LLM Security
von: Li, Shen, et al.
Veröffentlicht: (2024)
von: Li, Shen, et al.
Veröffentlicht: (2024)
A Framework for Formalizing LLM Agent Security
von: Siu, Vincent, et al.
Veröffentlicht: (2026)
von: Siu, Vincent, et al.
Veröffentlicht: (2026)
RTLMarker: Protecting LLM-Generated RTL Copyright via a Hardware Watermarking Framework
von: Wang, Kun, et al.
Veröffentlicht: (2025)
von: Wang, Kun, et al.
Veröffentlicht: (2025)
QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety
von: Lee, Taegyeong, et al.
Veröffentlicht: (2025)
von: Lee, Taegyeong, et al.
Veröffentlicht: (2025)
Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis
von: Topol, Zvi
Veröffentlicht: (2026)
von: Topol, Zvi
Veröffentlicht: (2026)
LLM-CSEC: Empirical Evaluation of Security in C/C++ Code Generated by Large Language Models
von: Shahid, Muhammad Usman, et al.
Veröffentlicht: (2025)
von: Shahid, Muhammad Usman, et al.
Veröffentlicht: (2025)
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
von: Zheng, Baolin, et al.
Veröffentlicht: (2025)
von: Zheng, Baolin, et al.
Veröffentlicht: (2025)
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
von: Hossain, Saad, et al.
Veröffentlicht: (2026)
von: Hossain, Saad, et al.
Veröffentlicht: (2026)
Fast Proxies for LLM Robustness Evaluation
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
von: Chen, Jizhou, et al.
Veröffentlicht: (2025)
von: Chen, Jizhou, et al.
Veröffentlicht: (2025)
SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration
von: Pan, Yu, et al.
Veröffentlicht: (2026)
von: Pan, Yu, et al.
Veröffentlicht: (2026)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
von: Zhang, Ruyi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruyi, et al.
Veröffentlicht: (2026)
LLM Access Shield: Domain-Specific LLM Framework for Privacy Policy Compliance
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
An LLM Framework For Cryptography Over Chat Channels
von: Gligoroski, Danilo, et al.
Veröffentlicht: (2025)
von: Gligoroski, Danilo, et al.
Veröffentlicht: (2025)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
von: Yang, Chenglin
Veröffentlicht: (2026)
von: Yang, Chenglin
Veröffentlicht: (2026)
IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computation
von: Guo, Yanpei, et al.
Veröffentlicht: (2026)
von: Guo, Yanpei, et al.
Veröffentlicht: (2026)
A Framework for Evaluating Emerging Cyberattack Capabilities of AI
von: Rodriguez, Mikel, et al.
Veröffentlicht: (2025)
von: Rodriguez, Mikel, et al.
Veröffentlicht: (2025)
FAST-IDS: A Fast Two-Stage Intrusion Detection System with Hybrid Compression for Real-Time Threat Detection in Connected and Autonomous Vehicles
von: S, Devika, et al.
Veröffentlicht: (2025)
von: S, Devika, et al.
Veröffentlicht: (2025)
ThreatGPT: An Agentic AI Framework for Enhancing Public Safety through Threat Modeling
von: Zisad, Sharif Noor, et al.
Veröffentlicht: (2025)
von: Zisad, Sharif Noor, et al.
Veröffentlicht: (2025)
CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning
von: ElZemity, Adel, et al.
Veröffentlicht: (2025)
von: ElZemity, Adel, et al.
Veröffentlicht: (2025)
A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
von: Wei, Qianshan, et al.
Veröffentlicht: (2025)
von: Wei, Qianshan, et al.
Veröffentlicht: (2025)
MALCDF: A Distributed Multi-Agent LLM Framework for Real-Time Cyber
von: Bhardwaj, Arth, et al.
Veröffentlicht: (2025)
von: Bhardwaj, Arth, et al.
Veröffentlicht: (2025)
A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense
von: Zhai, Keke
Veröffentlicht: (2024)
von: Zhai, Keke
Veröffentlicht: (2024)
Enhancing Security in LLM Applications: A Performance Evaluation of Early Detection Systems
von: Gakh, Valerii, et al.
Veröffentlicht: (2025)
von: Gakh, Valerii, et al.
Veröffentlicht: (2025)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
von: Vero, Mark, et al.
Veröffentlicht: (2026)
von: Vero, Mark, et al.
Veröffentlicht: (2026)
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
von: Tong, Haibo, et al.
Veröffentlicht: (2026)
von: Tong, Haibo, et al.
Veröffentlicht: (2026)
Efficient LLM Safety Evaluation through Multi-Agent Debate
von: Lin, Dachuan, et al.
Veröffentlicht: (2025)
von: Lin, Dachuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLM Safety for Children
von: Rath, Prasanjit, et al.
Veröffentlicht: (2025) -
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024) -
LLM-Safety Evaluations Lack Robustness
von: Beyer, Tim, et al.
Veröffentlicht: (2025) -
Security of and by Generative AI platforms
von: Hayagreevan, Hari, et al.
Veröffentlicht: (2024) -
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
von: Ji, Zimo, et al.
Veröffentlicht: (2025)