Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Collu, Matteo Gioele, Janssen-Groesbeek, Tom, Koffas, Stefanos, Conti, Mauro, Picek, Stjepan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025)
by: Chakraborty, Amit, et al.
Published: (2025)
ARMOUR US: Android Runtime Zero-permission Sensor Usage Monitoring from User Space
by: Long, Yan, et al.
Published: (2025)
by: Long, Yan, et al.
Published: (2025)
Memory Forensics Techniques for Automated Detection and Analysis of Go Malware
by: Ali, Hala, et al.
Published: (2026)
by: Ali, Hala, et al.
Published: (2026)
The Discovery, Disclosure, and Investigation of CVE-2024-25825
by: Chasens, Hunter
Published: (2025)
by: Chasens, Hunter
Published: (2025)
Post-quantum Federated Learning: Secure And Scalable Threat Intelligence For Collaborative Cyber Defense
by: Nayak, Prabhudarshi, et al.
Published: (2026)
by: Nayak, Prabhudarshi, et al.
Published: (2026)
Silent Consent, Persistent Risk: Android Permission Groups and Custom Permissions
by: Akanji, Olawale Amos, et al.
Published: (2026)
by: Akanji, Olawale Amos, et al.
Published: (2026)
Dynamic Vulnerability Patching for Heterogeneous Embedded Systems Using Stack Frame Reconstruction
by: Zhou, Ming, et al.
Published: (2025)
by: Zhou, Ming, et al.
Published: (2025)
vCause: Efficient and Verifiable Causality Analysis for Cloud-based Endpoint Auditing
by: Song, Qiyang, et al.
Published: (2026)
by: Song, Qiyang, et al.
Published: (2026)
DiMEx: Breaking the Cold Start Barrier in Data-Free Model Extraction via Latent Diffusion Priors
by: Thesia, Yash, et al.
Published: (2026)
by: Thesia, Yash, et al.
Published: (2026)
ReconXF: Graph Reconstruction Attack via Public Feature Explanations on Privatized Node Features and Labels
by: Sahoo, Rishi Raj, et al.
Published: (2025)
by: Sahoo, Rishi Raj, et al.
Published: (2025)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
by: Yu, Hao, et al.
Published: (2026)
by: Yu, Hao, et al.
Published: (2026)
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
by: Hoang, Tien Dat
Published: (2025)
by: Hoang, Tien Dat
Published: (2025)
The Need for Standardized Evidence Sampling in CMMC Assessments: A Survey-Based Analysis of Assessor Practices
by: Therrien, Logan, et al.
Published: (2026)
by: Therrien, Logan, et al.
Published: (2026)
Fine-tuning RoBERTa for CVE-to-CWE Classification: A 125M Parameter Model Competitive with LLMs
by: Mosievskiy, Nikita
Published: (2026)
by: Mosievskiy, Nikita
Published: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
SAFE-SiP: Secure Authentication Framework for System-in-Package Using Multi-party Computation
by: Tashdid, Ishraq, et al.
Published: (2025)
by: Tashdid, Ishraq, et al.
Published: (2025)
Toward Individual Fairness Without Centralized Data: Selective Counterfactual Consistency for Vertical Federated Learning
by: Wasif, Dawood, et al.
Published: (2026)
by: Wasif, Dawood, et al.
Published: (2026)
Static Attribution of Android Residential Proxy Malware Using Graph Kernels
by: Clark, Peter, et al.
Published: (2026)
by: Clark, Peter, et al.
Published: (2026)
ZTD$_{JAVA}$: Mitigating Software Supply Chain Vulnerabilities via Zero-Trust Dependencies
by: Amusuo, Paschal C., et al.
Published: (2023)
by: Amusuo, Paschal C., et al.
Published: (2023)
Fool Me If You Can: On the Robustness of Binary Code Similarity Detection Models against Semantics-preserving Transformations
by: Uhm, Jiyong, et al.
Published: (2026)
by: Uhm, Jiyong, et al.
Published: (2026)
UnPII: Unlearning Personally Identifiable Information with Quantifiable Exposure Risk
by: Jeon, Intae, et al.
Published: (2026)
by: Jeon, Intae, et al.
Published: (2026)
SpyChain: Multi-Vector Supply Chain Attacks on Small Satellite Systems
by: Vanlyssel, Jack, et al.
Published: (2025)
by: Vanlyssel, Jack, et al.
Published: (2025)
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
by: Nayak, Prabhudarshi, et al.
Published: (2026)
by: Nayak, Prabhudarshi, et al.
Published: (2026)
Operationalizing Cybersecurity Governance for Mitigation Planning with Attack-Path Modeling and Reinforcement Learning
by: Huff, Philip, et al.
Published: (2026)
by: Huff, Philip, et al.
Published: (2026)
Towards Low-Latency and Adaptive Ransomware Detection Using Contrastive Learning
by: Pan, Zhixin, et al.
Published: (2025)
by: Pan, Zhixin, et al.
Published: (2025)
Density-aware Sample-specific Attack
by: Wang, Qiyuan, et al.
Published: (2026)
by: Wang, Qiyuan, et al.
Published: (2026)
Risk-Calibrated Bayesian Streaming Intrusion Detection with SRE-Aligned Decisions
by: Youssef, Michel
Published: (2025)
by: Youssef, Michel
Published: (2025)
Broken Object Level Authorization in the Wild: An Empirical Taxonomy from 100+ Bug Bounty Disclosures
by: Kaur, Bandana
Published: (2026)
by: Kaur, Bandana
Published: (2026)
SecureBank: A Financially-Aware Zero Trust Architecture for High-Assurance Banking Systems
by: Biao, Paulo Fernandes
Published: (2025)
by: Biao, Paulo Fernandes
Published: (2025)
CryptoGuard: Lightweight Hybrid Detection and Response to Host-based Cryptojackers in Linux Cloud Environments
by: Park, Gyeonghoon, et al.
Published: (2025)
by: Park, Gyeonghoon, et al.
Published: (2025)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
by: Dawson, Ads, et al.
Published: (2025)
by: Dawson, Ads, et al.
Published: (2025)
Privacy-Enhancing Encryption in Data Sharing: A Survey on Security, Performance and Functionality
by: Lv, Yongyang, et al.
Published: (2026)
by: Lv, Yongyang, et al.
Published: (2026)
Witnessd: Proof-of-process via Adversarial Collapse
by: Condrey, David
Published: (2026)
by: Condrey, David
Published: (2026)
The Automation Advantage in AI Red Teaming
by: Mulla, Rob, et al.
Published: (2025)
by: Mulla, Rob, et al.
Published: (2025)
A High-Recall Cost-Sensitive Machine Learning Framework for Real-Time Online Banking Transaction Fraud Detection
by: R., Karthikeyan V., et al.
Published: (2026)
by: R., Karthikeyan V., et al.
Published: (2026)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
Non-Adaptive Adversarial Face Generation
by: Kim, Sunpill, et al.
Published: (2025)
by: Kim, Sunpill, et al.
Published: (2025)
Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI
by: Pasupuleti, Vinil, et al.
Published: (2026)
by: Pasupuleti, Vinil, et al.
Published: (2026)
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
by: Doda, Shravan
Published: (2026)
by: Doda, Shravan
Published: (2026)
Sensitivity Uncertainty Alignment in Large Language Models
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
Similar Items
-
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025) -
ARMOUR US: Android Runtime Zero-permission Sensor Usage Monitoring from User Space
by: Long, Yan, et al.
Published: (2025) -
Memory Forensics Techniques for Automated Detection and Analysis of Go Malware
by: Ali, Hala, et al.
Published: (2026) -
The Discovery, Disclosure, and Investigation of CVE-2024-25825
by: Chasens, Hunter
Published: (2025) -
Post-quantum Federated Learning: Secure And Scalable Threat Intelligence For Collaborative Cyber Defense
by: Nayak, Prabhudarshi, et al.
Published: (2026)