Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Halder, Subho, Saxena, Siddharth, Shrish, Kashinath Kadaba, M, Thiyagarajan |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Expert Interchange Standard: Enabling Dynamic Expert Management in Mixture-of-Experts Language Models
par: Kashinath, Kadaba Shrish
Publié: (2026)
par: Kashinath, Kadaba Shrish
Publié: (2026)
Capturing the security expert knowledge in feature selection for web application attack detection
par: Riverol, Amanda, et autres
Publié: (2024)
par: Riverol, Amanda, et autres
Publié: (2024)
CIPHER: Cryptographic Insecurity Profiling via Hybrid Evaluation of Responses
par: Manolov, Max, et autres
Publié: (2026)
par: Manolov, Max, et autres
Publié: (2026)
Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models
par: Nakka, Kalyan, et autres
Publié: (2024)
par: Nakka, Kalyan, et autres
Publié: (2024)
LLM-Safety Evaluations Lack Robustness
par: Beyer, Tim, et autres
Publié: (2025)
par: Beyer, Tim, et autres
Publié: (2025)
Fast Proxies for LLM Robustness Evaluation
par: Beyer, Tim, et autres
Publié: (2025)
par: Beyer, Tim, et autres
Publié: (2025)
RedTeamLLM: an Agentic AI framework for offensive security
par: Challita, Brian, et autres
Publié: (2025)
par: Challita, Brian, et autres
Publié: (2025)
NFC based inventory control system for secure and efficient communication
par: Iqbal, Razi, et autres
Publié: (2026)
par: Iqbal, Razi, et autres
Publié: (2026)
DP-FlogTinyLLM: Differentially private federated log anomaly detection using Tiny LLMs
par: Thompson, Isaiah, et autres
Publié: (2026)
par: Thompson, Isaiah, et autres
Publié: (2026)
AI-Driven IRM: Transforming insider risk management with adaptive scoring and LLM-based threat detection
par: Koli, Lokesh, et autres
Publié: (2025)
par: Koli, Lokesh, et autres
Publié: (2025)
Epistemic Bias Injection: Biasing LLMs via Selective Context Retrieval
par: Wu, Hao, et autres
Publié: (2025)
par: Wu, Hao, et autres
Publié: (2025)
SAGE: A Generic Framework for LLM Safety Evaluation
par: Jindal, Madhur, et autres
Publié: (2025)
par: Jindal, Madhur, et autres
Publié: (2025)
Robustness of LLM-enabled vehicle trajectory prediction under data security threats
par: Wang, Feilong, et autres
Publié: (2025)
par: Wang, Feilong, et autres
Publié: (2025)
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
par: Gao, Jiaxin, et autres
Publié: (2025)
par: Gao, Jiaxin, et autres
Publié: (2025)
LlamaFirewall: An open source guardrail system for building secure AI agents
par: Chennabasappa, Sahana, et autres
Publié: (2025)
par: Chennabasappa, Sahana, et autres
Publié: (2025)
LLMs on support of privacy and security of mobile apps: state of the art and research directions
par: Nguyen, Tran Thanh Lam, et autres
Publié: (2025)
par: Nguyen, Tran Thanh Lam, et autres
Publié: (2025)
COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
par: Wang, Junyu, et autres
Publié: (2025)
par: Wang, Junyu, et autres
Publié: (2025)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
par: Che, Zora, et autres
Publié: (2025)
par: Che, Zora, et autres
Publié: (2025)
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
par: Ji, Zimo, et autres
Publié: (2025)
par: Ji, Zimo, et autres
Publié: (2025)
Doppelganger Method: Breaking Role Consistency in LLM Agent via Prompt-based Transferable Adversarial Attack
par: Kang, Daewon, et autres
Publié: (2025)
par: Kang, Daewon, et autres
Publié: (2025)
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks
par: Zhao, Jin, et autres
Publié: (2026)
par: Zhao, Jin, et autres
Publié: (2026)
Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks
par: Sahu, Anubhab, et autres
Publié: (2026)
par: Sahu, Anubhab, et autres
Publié: (2026)
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
par: Wang, Shouju, et autres
Publié: (2025)
par: Wang, Shouju, et autres
Publié: (2025)
Enhancing Security in LLM Applications: A Performance Evaluation of Early Detection Systems
par: Gakh, Valerii, et autres
Publié: (2025)
par: Gakh, Valerii, et autres
Publié: (2025)
ZeroDayBench: Evaluating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense
par: Lau, Nancy, et autres
Publié: (2026)
par: Lau, Nancy, et autres
Publié: (2026)
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents
par: Liu, Shi, et autres
Publié: (2026)
par: Liu, Shi, et autres
Publié: (2026)
Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks
par: De Stefano, Gianluca, et autres
Publié: (2024)
par: De Stefano, Gianluca, et autres
Publié: (2024)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
par: Fu, Yuchuan, et autres
Publié: (2025)
par: Fu, Yuchuan, et autres
Publié: (2025)
LLM-CSEC: Empirical Evaluation of Security in C/C++ Code Generated by Large Language Models
par: Shahid, Muhammad Usman, et autres
Publié: (2025)
par: Shahid, Muhammad Usman, et autres
Publié: (2025)
Design and implementation of a distributed security threat detection system integrating federated learning and multimodal LLM
par: Wang, Yuqing, et autres
Publié: (2025)
par: Wang, Yuqing, et autres
Publié: (2025)
A Scalable Multi-GPU Framework for Encrypted Large-Model Inference
par: Jayashankar, Siddharth, et autres
Publié: (2025)
par: Jayashankar, Siddharth, et autres
Publié: (2025)
TADP-RME: A Trust-Adaptive Differential Privacy Framework for Enhancing Reliability of Data-Driven Systems
par: Halder, Labani, et autres
Publié: (2026)
par: Halder, Labani, et autres
Publié: (2026)
GraMFedDHAR: Graph Based Multimodal Differentially Private Federated HAR
par: Halder, Labani, et autres
Publié: (2025)
par: Halder, Labani, et autres
Publié: (2025)
Network evasion detection with Bi-LSTM model
par: Chen, Kehua, et autres
Publié: (2025)
par: Chen, Kehua, et autres
Publié: (2025)
Membership Inference Attacks fueled by Few-Short Learning to detect privacy leakage tackling data integrity
par: Jiménez-López, Daniel, et autres
Publié: (2025)
par: Jiménez-López, Daniel, et autres
Publié: (2025)
sudo rm -rf agentic_security
par: Lee, Sejin, et autres
Publié: (2025)
par: Lee, Sejin, et autres
Publié: (2025)
ActDroid: An active learning framework for Android malware detection
par: Muzaffar, Ali, et autres
Publié: (2024)
par: Muzaffar, Ali, et autres
Publié: (2024)
Anomaly detection in network flows using unsupervised online machine learning
par: Miguel-Diez, Alberto, et autres
Publié: (2025)
par: Miguel-Diez, Alberto, et autres
Publié: (2025)
The importance of the clustering model to detect new types of intrusion in data traffic
par: Abd, Noor Saud, et autres
Publié: (2024)
par: Abd, Noor Saud, et autres
Publié: (2024)
COPS: A Compact On-device Pipeline for real-time Smishing detection
par: S, Harichandana B S, et autres
Publié: (2024)
par: S, Harichandana B S, et autres
Publié: (2024)
Documents similaires
-
The Expert Interchange Standard: Enabling Dynamic Expert Management in Mixture-of-Experts Language Models
par: Kashinath, Kadaba Shrish
Publié: (2026) -
Capturing the security expert knowledge in feature selection for web application attack detection
par: Riverol, Amanda, et autres
Publié: (2024) -
CIPHER: Cryptographic Insecurity Profiling via Hybrid Evaluation of Responses
par: Manolov, Max, et autres
Publié: (2026) -
Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models
par: Nakka, Kalyan, et autres
Publié: (2024) -
LLM-Safety Evaluations Lack Robustness
par: Beyer, Tim, et autres
Publié: (2025)