Amplifying Training Data Exposure through Fine-Tuning with Pseudo-Labeled Memberships
Fuente:
arXiv
Salvato in:
| Autori principali: | Oh, Myung Gyo, Ahn, Hong Eun, Park, Leo Hyun, Kwon, Taekyoung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adversarial Feature Alignment: Balancing Robustness and Accuracy in Deep Learning via Adversarial Training
di: Park, Leo Hyun, et al.
Pubblicazione: (2024)
di: Park, Leo Hyun, et al.
Pubblicazione: (2024)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
di: Gu, Yongtong, et al.
Pubblicazione: (2026)
di: Gu, Yongtong, et al.
Pubblicazione: (2026)
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
di: Morales, Jaime, et al.
Pubblicazione: (2026)
di: Morales, Jaime, et al.
Pubblicazione: (2026)
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts
di: Young, Richard J., et al.
Pubblicazione: (2026)
di: Young, Richard J., et al.
Pubblicazione: (2026)
Send to which account? Evaluation of an LLM-based Scambaiting System
di: Siadati, Hossein, et al.
Pubblicazione: (2025)
di: Siadati, Hossein, et al.
Pubblicazione: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
di: Chona, Alankrit, et al.
Pubblicazione: (2026)
di: Chona, Alankrit, et al.
Pubblicazione: (2026)
Measuring Harmfulness of Computer-Using Agents
di: Tian, Aaron Xuxiang, et al.
Pubblicazione: (2025)
di: Tian, Aaron Xuxiang, et al.
Pubblicazione: (2025)
Countermind: A Multi-Layered Security Architecture for Large Language Models
di: Schwarz, Dominik
Pubblicazione: (2025)
di: Schwarz, Dominik
Pubblicazione: (2025)
Semantic Superiority vs. Forensic Efficiency: A Comparative Analysis of Deep Learning and Psycholinguistics for Business Email Compromise Detection
di: Adjei, Yaw Osei, et al.
Pubblicazione: (2025)
di: Adjei, Yaw Osei, et al.
Pubblicazione: (2025)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
di: Dawson, Ads, et al.
Pubblicazione: (2025)
di: Dawson, Ads, et al.
Pubblicazione: (2025)
The Automation Advantage in AI Red Teaming
di: Mulla, Rob, et al.
Pubblicazione: (2025)
di: Mulla, Rob, et al.
Pubblicazione: (2025)
Towards Modeling Cybersecurity Behavior of Humans in Organizations
di: Kürtz, Klaas Ole
Pubblicazione: (2026)
di: Kürtz, Klaas Ole
Pubblicazione: (2026)
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
di: Wang, Haochuan Kevin, et al.
Pubblicazione: (2026)
di: Wang, Haochuan Kevin, et al.
Pubblicazione: (2026)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
di: Othman, Refat
Pubblicazione: (2026)
di: Othman, Refat
Pubblicazione: (2026)
CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments
di: Keppler, Gustav, et al.
Pubblicazione: (2026)
di: Keppler, Gustav, et al.
Pubblicazione: (2026)
Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models
di: Syed, Mohammed Sameer, et al.
Pubblicazione: (2026)
di: Syed, Mohammed Sameer, et al.
Pubblicazione: (2026)
SALLIE: Safeguarding Against Latent Language & Image Exploits
di: Azov, Guy, et al.
Pubblicazione: (2026)
di: Azov, Guy, et al.
Pubblicazione: (2026)
SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
di: Adebimpe, Adetayo, et al.
Pubblicazione: (2025)
di: Adebimpe, Adetayo, et al.
Pubblicazione: (2025)
Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests
di: Young, Richard J., et al.
Pubblicazione: (2026)
di: Young, Richard J., et al.
Pubblicazione: (2026)
Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management
di: Arora, Sunil, et al.
Pubblicazione: (2025)
di: Arora, Sunil, et al.
Pubblicazione: (2025)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
di: Dang, Kieu, et al.
Pubblicazione: (2025)
di: Dang, Kieu, et al.
Pubblicazione: (2025)
Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
di: Usman, Rana Muhammad
Pubblicazione: (2026)
di: Usman, Rana Muhammad
Pubblicazione: (2026)
Retrieval Augmented Classification for Confidential Documents
di: Chang, Yeseul E., et al.
Pubblicazione: (2026)
di: Chang, Yeseul E., et al.
Pubblicazione: (2026)
Powerful Training-Free Membership Inference Against Autoregressive Language Models
di: Ilić, David, et al.
Pubblicazione: (2026)
di: Ilić, David, et al.
Pubblicazione: (2026)
Adaptive Defense Orchestration for RAG: A Sentinel-Strategist Architecture against Multi-Vector Attacks
di: Pallerla, Pranav, et al.
Pubblicazione: (2026)
di: Pallerla, Pranav, et al.
Pubblicazione: (2026)
Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility
di: Engelberg, Gal, et al.
Pubblicazione: (2026)
di: Engelberg, Gal, et al.
Pubblicazione: (2026)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
di: Grofsky, Matthew
Pubblicazione: (2025)
di: Grofsky, Matthew
Pubblicazione: (2025)
Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study
di: Khatiwala, Jeel Piyushkumar, et al.
Pubblicazione: (2026)
di: Khatiwala, Jeel Piyushkumar, et al.
Pubblicazione: (2026)
Multilingual AI-Driven Password Strength Estimation with Similarity-Based Detection
di: Palaniappan, Nikitha M., et al.
Pubblicazione: (2026)
di: Palaniappan, Nikitha M., et al.
Pubblicazione: (2026)
Towards Agentic Investigation of Security Alerts
di: Eilertsen, Even, et al.
Pubblicazione: (2026)
di: Eilertsen, Even, et al.
Pubblicazione: (2026)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
di: Zhang, Tian, et al.
Pubblicazione: (2026)
di: Zhang, Tian, et al.
Pubblicazione: (2026)
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
di: Chen, Renmiao, et al.
Pubblicazione: (2025)
di: Chen, Renmiao, et al.
Pubblicazione: (2025)
Whisper Leak: a side-channel attack on Large Language Models
di: McDonald, Geoff, et al.
Pubblicazione: (2025)
di: McDonald, Geoff, et al.
Pubblicazione: (2025)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
di: Ge, Yuxu
Pubblicazione: (2026)
di: Ge, Yuxu
Pubblicazione: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
di: Young, Richard J., et al.
Pubblicazione: (2026)
di: Young, Richard J., et al.
Pubblicazione: (2026)
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
di: Lv, Bo, et al.
Pubblicazione: (2026)
di: Lv, Bo, et al.
Pubblicazione: (2026)
LLM Detectors Still Fall Short of Real World: Case of LLM-Generated Short News-Like Posts
di: Gameiro, Henrique Da Silva, et al.
Pubblicazione: (2024)
di: Gameiro, Henrique Da Silva, et al.
Pubblicazione: (2024)
Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models
di: Ntais, Pavlos
Pubblicazione: (2025)
di: Ntais, Pavlos
Pubblicazione: (2025)
Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks
di: Hu, Saisai
Pubblicazione: (2026)
di: Hu, Saisai
Pubblicazione: (2026)
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
di: Lelle, Travis
Pubblicazione: (2026)
di: Lelle, Travis
Pubblicazione: (2026)
Documenti analoghi
-
Adversarial Feature Alignment: Balancing Robustness and Accuracy in Deep Learning via Adversarial Training
di: Park, Leo Hyun, et al.
Pubblicazione: (2024) -
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
di: Gu, Yongtong, et al.
Pubblicazione: (2026) -
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
di: Morales, Jaime, et al.
Pubblicazione: (2026) -
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts
di: Young, Richard J., et al.
Pubblicazione: (2026) -
Send to which account? Evaluation of an LLM-based Scambaiting System
di: Siadati, Hossein, et al.
Pubblicazione: (2025)