S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Yuan, Xiaohan, Li, Jinfeng, Wang, Dongxia, Chen, Yuefeng, Mao, Xiaofeng, Huang, Longtao, Chen, Jialuo, Xue, Hui, Liu, Xiaoxia, Wang, Wenhai, Ren, Kui, Wang, Jingyi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
di: Fu, Yuchuan, et al.
Pubblicazione: (2025)
di: Fu, Yuchuan, et al.
Pubblicazione: (2025)
FoolSDEdit: Deceptively Steering Your Edits Towards Targeted Attribute-aware Distribution
di: Zhou, Qi, et al.
Pubblicazione: (2024)
di: Zhou, Qi, et al.
Pubblicazione: (2024)
MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
di: Chen, Tailun, et al.
Pubblicazione: (2025)
di: Chen, Tailun, et al.
Pubblicazione: (2025)
CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity
di: Yu, Zhengmin, et al.
Pubblicazione: (2024)
di: Yu, Zhengmin, et al.
Pubblicazione: (2024)
Toward Unbiased Multiple-Target Fuzzing with Path Diversity
di: Rong, Huanyao, et al.
Pubblicazione: (2023)
di: Rong, Huanyao, et al.
Pubblicazione: (2023)
Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
di: Chen, Xiangxiang, et al.
Pubblicazione: (2025)
di: Chen, Xiangxiang, et al.
Pubblicazione: (2025)
Autonomous LLM Agent Worms: Cross-Platform Propagation, Automated Discovery and Temporal Re-Entry Defense
di: Zha, Mingming, et al.
Pubblicazione: (2026)
di: Zha, Mingming, et al.
Pubblicazione: (2026)
Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models
di: Qi, Weiwei, et al.
Pubblicazione: (2026)
di: Qi, Weiwei, et al.
Pubblicazione: (2026)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
di: Li, Tianhao, et al.
Pubblicazione: (2024)
di: Li, Tianhao, et al.
Pubblicazione: (2024)
PRUNE: A Patching Based Repair Framework for Certifiable Unlearning of Neural Networks
di: Li, Xuran, et al.
Pubblicazione: (2025)
di: Li, Xuran, et al.
Pubblicazione: (2025)
SoK: Towards Effective Automated Vulnerability Repair
di: Li, Ying, et al.
Pubblicazione: (2025)
di: Li, Ying, et al.
Pubblicazione: (2025)
VulEval: Towards Repository-Level Evaluation of Software Vulnerability Detection
di: Wen, Xin-Cheng, et al.
Pubblicazione: (2024)
di: Wen, Xin-Cheng, et al.
Pubblicazione: (2024)
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
di: Chen, Tianyu, et al.
Pubblicazione: (2026)
di: Chen, Tianyu, et al.
Pubblicazione: (2026)
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
di: Zhang, Jiawen, et al.
Pubblicazione: (2025)
di: Zhang, Jiawen, et al.
Pubblicazione: (2025)
From Sands to Mansions: Towards Automated Cyberattack Emulation with Classical Planning and Large Language Models
di: Wang, Lingzhi, et al.
Pubblicazione: (2024)
di: Wang, Lingzhi, et al.
Pubblicazione: (2024)
Generalized Security-Preserving Refinement for Concurrent Systems
di: Sun, Huan, et al.
Pubblicazione: (2025)
di: Sun, Huan, et al.
Pubblicazione: (2025)
PortGPT: Towards Automated Backporting Using Large Language Models
di: Li, Zhaoyang, et al.
Pubblicazione: (2025)
di: Li, Zhaoyang, et al.
Pubblicazione: (2025)
Mitigating Data Poisoning Attacks to Local Differential Privacy
di: Li, Xiaolin, et al.
Pubblicazione: (2025)
di: Li, Xiaolin, et al.
Pubblicazione: (2025)
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
di: Zhang, Wenhui, et al.
Pubblicazione: (2025)
di: Zhang, Wenhui, et al.
Pubblicazione: (2025)
Non-intrusive and Unconstrained Keystroke Inference in VR Platforms via Infrared Side Channel
di: Ni, Tao, et al.
Pubblicazione: (2024)
di: Ni, Tao, et al.
Pubblicazione: (2024)
A Certified Robust Watermark For Large Language Models
di: Feng, Xianheng, et al.
Pubblicazione: (2024)
di: Feng, Xianheng, et al.
Pubblicazione: (2024)
KryptoPilot: An Open-World Knowledge-Augmented LLM Agent for Automated Cryptographic Exploitation
di: Liu, Xiaonan, et al.
Pubblicazione: (2026)
di: Liu, Xiaonan, et al.
Pubblicazione: (2026)
AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models
di: Wang, Jackson
Pubblicazione: (2026)
di: Wang, Jackson
Pubblicazione: (2026)
Towards Automated Discovery of Asymmetric Mempool DoS in Blockchains
di: Wang, Yibo, et al.
Pubblicazione: (2023)
di: Wang, Yibo, et al.
Pubblicazione: (2023)
CyberThreat-Eval: Can Large Language Models Automate Real-World Threat Research?
di: Chen, Xiangsen, et al.
Pubblicazione: (2026)
di: Chen, Xiangsen, et al.
Pubblicazione: (2026)
Towards Imperceptible Adversarial Defense: A Gradient-Driven Shield against Facial Manipulations
di: Li, Yue, et al.
Pubblicazione: (2025)
di: Li, Yue, et al.
Pubblicazione: (2025)
LoRA-Key: User-Centric LoRA Watermarking for Text-to-Image Diffusion Models
di: Wang, Yaopeng, et al.
Pubblicazione: (2026)
di: Wang, Yaopeng, et al.
Pubblicazione: (2026)
WALLETRADAR: Towards Automating the Detection of Vulnerabilities in Browser-based Cryptocurrency Wallets
di: Xia, Pengcheng, et al.
Pubblicazione: (2024)
di: Xia, Pengcheng, et al.
Pubblicazione: (2024)
PT-Mark: Invisible Watermarking for Text-to-image Diffusion Models via Semantic-aware Pivotal Tuning
di: Wang, Yaopeng, et al.
Pubblicazione: (2025)
di: Wang, Yaopeng, et al.
Pubblicazione: (2025)
Safeguarding LLM Embeddings in End-Cloud Collaboration via Entropy-Driven Perturbation
di: Jin, Shuaifan, et al.
Pubblicazione: (2025)
di: Jin, Shuaifan, et al.
Pubblicazione: (2025)
AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
di: Liang, Jiacheng, et al.
Pubblicazione: (2025)
di: Liang, Jiacheng, et al.
Pubblicazione: (2025)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
di: Liu, Songyang, et al.
Pubblicazione: (2026)
di: Liu, Songyang, et al.
Pubblicazione: (2026)
LoopTrap: Termination Poisoning Attacks on LLM Agents
di: Xu, Huiyu, et al.
Pubblicazione: (2026)
di: Xu, Huiyu, et al.
Pubblicazione: (2026)
PentestAgent: Incorporating LLM Agents to Automated Penetration Testing
di: Shen, Xiangmin, et al.
Pubblicazione: (2024)
di: Shen, Xiangmin, et al.
Pubblicazione: (2024)
LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
di: Li, Songze, et al.
Pubblicazione: (2025)
di: Li, Songze, et al.
Pubblicazione: (2025)
Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads
di: Wu, Jinman, et al.
Pubblicazione: (2026)
di: Wu, Jinman, et al.
Pubblicazione: (2026)
KV-Auditor: Auditing Local Differential Privacy for Correlated Key-Value Estimation
di: Xu, Jingnan, et al.
Pubblicazione: (2025)
di: Xu, Jingnan, et al.
Pubblicazione: (2025)
SWAT: A System-Wide Approach to Tunable Leakage Mitigation in Encrypted Data Stores
di: Zheng, Leqian, et al.
Pubblicazione: (2023)
di: Zheng, Leqian, et al.
Pubblicazione: (2023)
BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model
di: Lin, Weilin, et al.
Pubblicazione: (2025)
di: Lin, Weilin, et al.
Pubblicazione: (2025)
Cuckoo Attack: Stealthy and Persistent Attacks Against AI-IDE
di: Liu, Xinpeng, et al.
Pubblicazione: (2025)
di: Liu, Xinpeng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
di: Fu, Yuchuan, et al.
Pubblicazione: (2025) -
FoolSDEdit: Deceptively Steering Your Edits Towards Targeted Attribute-aware Distribution
di: Zhou, Qi, et al.
Pubblicazione: (2024) -
MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
di: Chen, Tailun, et al.
Pubblicazione: (2025) -
CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity
di: Yu, Zhengmin, et al.
Pubblicazione: (2024) -
Toward Unbiased Multiple-Target Fuzzing with Path Diversity
di: Rong, Huanyao, et al.
Pubblicazione: (2023)