Exploring Advanced Methodologies in Security Evaluation for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Jun, Zhang, Jiawei, Wang, Qi, Han, Weihong, Zhang, Yanchun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
by: Shao, Minghao, et al.
Published: (2024)
by: Shao, Minghao, et al.
Published: (2024)
LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
by: Ullah, Saad, et al.
Published: (2023)
by: Ullah, Saad, et al.
Published: (2023)
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
by: Han, Junxiao, et al.
Published: (2025)
by: Han, Junxiao, et al.
Published: (2025)
LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers
by: Okada, Hiroyuki, et al.
Published: (2026)
by: Okada, Hiroyuki, et al.
Published: (2026)
Secure Low-altitude Maritime Communications via Intelligent Jamming
by: Huang, Jiawei, et al.
Published: (2025)
by: Huang, Jiawei, et al.
Published: (2025)
Before You Hand Over the Wheel: Evaluating LLMs for Security Incident Analysis
by: Jajodia, Sourov, et al.
Published: (2026)
by: Jajodia, Sourov, et al.
Published: (2026)
Are LLMs Vulnerable to Preference-Undermining Attacks (PUA)? A Factorial Analysis Methodology for Diagnosing the Trade-off between Preference Alignment and Real-World Validity
by: An, Hongjun, et al.
Published: (2026)
by: An, Hongjun, et al.
Published: (2026)
FedSecurity: Benchmarking Attacks and Defenses in Federated Learning and Federated LLMs
by: Han, Shanshan, et al.
Published: (2023)
by: Han, Shanshan, et al.
Published: (2023)
LTRDetector: Exploring Long-Term Relationship for Advanced Persistent Threats Detection
by: Liu, Xiaoxiao, et al.
Published: (2024)
by: Liu, Xiaoxiao, et al.
Published: (2024)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Silent Guardians: Independent and Secure Decision Tree Evaluation Without Chatter
by: Li, Jinyuan, et al.
Published: (2026)
by: Li, Jinyuan, et al.
Published: (2026)
Security and Privacy on Generative Data in AIGC: A Survey
by: Wang, Tao, et al.
Published: (2023)
by: Wang, Tao, et al.
Published: (2023)
The Invitation Trap: Proactive Availability Backdoor in LLMs via Conversational Induction
by: Wang, He, et al.
Published: (2026)
by: Wang, He, et al.
Published: (2026)
Evaluation Methodologies in Software Protection Research
by: De Sutter, Bjorn, et al.
Published: (2023)
by: De Sutter, Bjorn, et al.
Published: (2023)
Exploring the Techniques of Information Security Certification
by: Chen, Abel C. H.
Published: (2023)
by: Chen, Abel C. H.
Published: (2023)
Exploring Power Side-Channel Challenges in Embedded Systems Security
by: Narimani, Pouya, et al.
Published: (2024)
by: Narimani, Pouya, et al.
Published: (2024)
Towards Compositional Generalization in LLMs for Smart Contract Security: A Case Study on Reentrancy Vulnerabilities
by: Zhou, Ying, et al.
Published: (2026)
by: Zhou, Ying, et al.
Published: (2026)
GIFDL: Generated Image Fluctuation Distortion Learning for Enhancing Steganographic Security
by: Wang, Xiangkun, et al.
Published: (2025)
by: Wang, Xiangkun, et al.
Published: (2025)
The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability
by: Xu, Zijie, et al.
Published: (2025)
by: Xu, Zijie, et al.
Published: (2025)
SafeClaw-R: Towards Safe and Secure Multi-Agent Personal Assistants
by: Wang, Haoyu, et al.
Published: (2026)
by: Wang, Haoyu, et al.
Published: (2026)
Advancing Honeywords for Real-World Authentication Security
by: Das, Sudiksha, et al.
Published: (2025)
by: Das, Sudiksha, et al.
Published: (2025)
Using LLMs for Tabletop Exercises within the Security Domain
by: Hays, Sam, et al.
Published: (2024)
by: Hays, Sam, et al.
Published: (2024)
Advanced Penetration Testing for Enhancing 5G Security
by: Smith-Haynes, Shari-Ann
Published: (2024)
by: Smith-Haynes, Shari-Ann
Published: (2024)
PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
by: Zhao, Jiawei, et al.
Published: (2025)
by: Zhao, Jiawei, et al.
Published: (2025)
Lightweight Yet Secure: Secure Scripting Language Generation via Lightweight LLMs
by: Zhang, Keyang, et al.
Published: (2026)
by: Zhang, Keyang, et al.
Published: (2026)
ARuleCon: Agentic Security Rule Conversion
by: Xu, Ming, et al.
Published: (2026)
by: Xu, Ming, et al.
Published: (2026)
RISecure-PUF: Multipurpose PUF-Driven Security Extensions with Lookaside Buffer in RISC-V
by: Chen, Chenghao, et al.
Published: (2024)
by: Chen, Chenghao, et al.
Published: (2024)
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI
by: Zhang, Zelin, et al.
Published: (2026)
by: Zhang, Zelin, et al.
Published: (2026)
SNNGX: Securing Spiking Neural Networks with Genetic XOR Encryption on RRAM-based Neuromorphic Accelerator
by: Wong, Kwunhang, et al.
Published: (2024)
by: Wong, Kwunhang, et al.
Published: (2024)
Can Developers rely on LLMs for Secure IaC Development?
by: Firouzi, Ehsan, et al.
Published: (2026)
by: Firouzi, Ehsan, et al.
Published: (2026)
Experimental Secure Multiparty Computation from Quantum Oblivious Transfer with Bit Commitment
by: Zhang, Kai-Yi, et al.
Published: (2024)
by: Zhang, Kai-Yi, et al.
Published: (2024)
Simple But Not Secure: An Empirical Security Analysis of Two-factor Authentication Systems
by: Wang, Zhi, et al.
Published: (2024)
by: Wang, Zhi, et al.
Published: (2024)
Securing the Open RAN Infrastructure: Exploring Vulnerabilities in Kubernetes Deployments
by: Klement, Felix, et al.
Published: (2024)
by: Klement, Felix, et al.
Published: (2024)
When FinTech Meets Privacy: Securing Financial LLMs with Differential Private Fine-Tuning
by: Zhu, Sichen, et al.
Published: (2025)
by: Zhu, Sichen, et al.
Published: (2025)
Generative AI for Secure and Privacy-Preserving Mobile Crowdsensing
by: Yang, Yaoqi, et al.
Published: (2024)
by: Yang, Yaoqi, et al.
Published: (2024)
Federated Learning for Cross-Domain Data Privacy: A Distributed Approach to Secure Collaboration
by: Zhang, Yiwei, et al.
Published: (2025)
by: Zhang, Yiwei, et al.
Published: (2025)
GuardPhish: Securing Open-Source LLMs from Phishing Abuse
by: Mishra, Rina, et al.
Published: (2026)
by: Mishra, Rina, et al.
Published: (2026)
Securing UAV Communications by Fusing Cross-Layer Fingerprints
by: Huang, Yong, et al.
Published: (2025)
by: Huang, Yong, et al.
Published: (2025)
A Comparative Evaluation of AI Agent Security Guardrails
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Fixing Security Vulnerabilities with AI in OSS-Fuzz
by: Zhang, Yuntong, et al.
Published: (2024)
by: Zhang, Yuntong, et al.
Published: (2024)
Similar Items
-
An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
by: Shao, Minghao, et al.
Published: (2024) -
LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
by: Ullah, Saad, et al.
Published: (2023) -
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
by: Han, Junxiao, et al.
Published: (2025) -
LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers
by: Okada, Hiroyuki, et al.
Published: (2026) -
Secure Low-altitude Maritime Communications via Intelligent Jamming
by: Huang, Jiawei, et al.
Published: (2025)