Provably Protecting Fine-Tuned LLMs from Training Data Extraction while Preserving Utility
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Segal, Tom, Shabtai, Asaf, Elovici, Yuval |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DOMBA: Double Model Balancing for Access-Controlled Language Models via Minimum-Bounded Aggregation
von: Segal, Tom, et al.
Veröffentlicht: (2024)
von: Segal, Tom, et al.
Veröffentlicht: (2024)
LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
von: Schwartz, Yuval, et al.
Veröffentlicht: (2024)
von: Schwartz, Yuval, et al.
Veröffentlicht: (2024)
Addressing Key Challenges of Adversarial Attacks and Defenses in the Tabular Domain: A Methodological Framework for Coherence and Consistency
von: Itzhakev, Yael, et al.
Veröffentlicht: (2024)
von: Itzhakev, Yael, et al.
Veröffentlicht: (2024)
RAPID: Robust APT Detection and Investigation Using Context-Aware Deep Learning
von: Amaru, Yonatan, et al.
Veröffentlicht: (2024)
von: Amaru, Yonatan, et al.
Veröffentlicht: (2024)
Real-World Adversarial Attacks on RF-Based Drone Detectors
von: Gazit, Omer, et al.
Veröffentlicht: (2025)
von: Gazit, Omer, et al.
Veröffentlicht: (2025)
QuantAttack: Exploiting Dynamic Quantization to Attack Vision Transformers
von: Baras, Amit, et al.
Veröffentlicht: (2023)
von: Baras, Amit, et al.
Veröffentlicht: (2023)
RuleGenie: SIEM Detection Rule Set Optimization
von: Shukla, Akansha, et al.
Veröffentlicht: (2025)
von: Shukla, Akansha, et al.
Veröffentlicht: (2025)
Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs
von: Rachmil, Oren, et al.
Veröffentlicht: (2025)
von: Rachmil, Oren, et al.
Veröffentlicht: (2025)
Rogue Cell: Adversarial Attack and Defense in Untrusted O-RAN Setup Exploiting the Traffic Steering xApp
von: Aizikovich, Eran, et al.
Veröffentlicht: (2025)
von: Aizikovich, Eran, et al.
Veröffentlicht: (2025)
Detection of Compromised Functions in a Serverless Cloud Environment
von: Lavi, Danielle, et al.
Veröffentlicht: (2024)
von: Lavi, Danielle, et al.
Veröffentlicht: (2024)
DIESEL -- Dynamic Inference-Guidance via Evasion of Semantic Embeddings in LLMs
von: Ganon, Ben, et al.
Veröffentlicht: (2024)
von: Ganon, Ben, et al.
Veröffentlicht: (2024)
DeSparsify: Adversarial Attack Against Token Sparsification Mechanisms in Vision Transformers
von: Yehezkel, Oryan, et al.
Veröffentlicht: (2024)
von: Yehezkel, Oryan, et al.
Veröffentlicht: (2024)
GenKubeSec: LLM-Based Kubernetes Misconfiguration Detection, Localization, Reasoning, and Remediation
von: Malul, Ehud, et al.
Veröffentlicht: (2024)
von: Malul, Ehud, et al.
Veröffentlicht: (2024)
AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior
von: Abaev, Nadya, et al.
Veröffentlicht: (2026)
von: Abaev, Nadya, et al.
Veröffentlicht: (2026)
KubeGuard: LLM-Assisted Kubernetes Hardening via Configuration Files and Runtime Logs Analysis
von: Cohen, Omri Sgan, et al.
Veröffentlicht: (2025)
von: Cohen, Omri Sgan, et al.
Veröffentlicht: (2025)
MIA-EPT: Membership Inference Attack via Error Prediction for Tabular Data
von: German, Eyal, et al.
Veröffentlicht: (2025)
von: German, Eyal, et al.
Veröffentlicht: (2025)
Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs
von: German, Eyal, et al.
Veröffentlicht: (2025)
von: German, Eyal, et al.
Veröffentlicht: (2025)
Tag&Tab: Pretraining Data Detection in Large Language Models Using Keyword-Based Membership Inference Attack
von: Antebi, Sagiv, et al.
Veröffentlicht: (2025)
von: Antebi, Sagiv, et al.
Veröffentlicht: (2025)
SoK: Cybersecurity Assessment of Humanoid Ecosystem
von: Surve, Priyanka Prakash, et al.
Veröffentlicht: (2025)
von: Surve, Priyanka Prakash, et al.
Veröffentlicht: (2025)
LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
von: German, Eyal, et al.
Veröffentlicht: (2025)
von: German, Eyal, et al.
Veröffentlicht: (2025)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
von: Noah, Amit Finkman, et al.
Veröffentlicht: (2024)
von: Noah, Amit Finkman, et al.
Veröffentlicht: (2024)
UEFI Memory Forensics: A Framework for UEFI Threat Analysis
von: Segal, Kalanit Suzan, et al.
Veröffentlicht: (2025)
von: Segal, Kalanit Suzan, et al.
Veröffentlicht: (2025)
PaniCar: Securing the Perception of Advanced Driving Assistance Systems Against Emergency Vehicle Lighting
von: Feldman, Elad, et al.
Veröffentlicht: (2025)
von: Feldman, Elad, et al.
Veröffentlicht: (2025)
FRAME : Comprehensive Risk Assessment Framework for Adversarial Machine Learning Threats
von: Shapira, Avishag, et al.
Veröffentlicht: (2025)
von: Shapira, Avishag, et al.
Veröffentlicht: (2025)
From Tool Orchestration to Code Execution: A Study of MCP Design Choices
von: Felendler, Yuval, et al.
Veröffentlicht: (2026)
von: Felendler, Yuval, et al.
Veröffentlicht: (2026)
Understanding and Preserving Safety in Fine-Tuned LLMs
von: Zhang, Jiawen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiawen, et al.
Veröffentlicht: (2026)
A Privacy Enhancing Technique to Evade Detection by Street Video Cameras Without Using Adversarial Accessories
von: Shams, Jacob, et al.
Veröffentlicht: (2025)
von: Shams, Jacob, et al.
Veröffentlicht: (2025)
ImpReSS: Implicit Recommender System for Support Conversations
von: Haller, Omri, et al.
Veröffentlicht: (2025)
von: Haller, Omri, et al.
Veröffentlicht: (2025)
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
von: Ran-Milo, Yuval
Veröffentlicht: (2026)
von: Ran-Milo, Yuval
Veröffentlicht: (2026)
Transferability Ranking of Adversarial Examples
von: Levy, Mosh, et al.
Veröffentlicht: (2022)
von: Levy, Mosh, et al.
Veröffentlicht: (2022)
Provable Imbalanced Point Clustering
von: Denisov, David, et al.
Veröffentlicht: (2024)
von: Denisov, David, et al.
Veröffentlicht: (2024)
Continual Fine-Tuning with Provably Accurate and Parameter-Free Task Retrieval
von: Le, Hang Thi-Thuy, et al.
Veröffentlicht: (2026)
von: Le, Hang Thi-Thuy, et al.
Veröffentlicht: (2026)
Rule-ATT&CK Mapper (RAM): Mapping SIEM Rules to TTPs Using LLMs
von: Wudali, Prasanna N., et al.
Veröffentlicht: (2025)
von: Wudali, Prasanna N., et al.
Veröffentlicht: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
von: Pan, Birong, et al.
Veröffentlicht: (2025)
von: Pan, Birong, et al.
Veröffentlicht: (2025)
Multi-Task LLM with LoRA Fine-Tuning for Automated Cancer Staging and Biomarker Extraction
von: Shao, Jiahao, et al.
Veröffentlicht: (2026)
von: Shao, Jiahao, et al.
Veröffentlicht: (2026)
FedMomentum: Preserving LoRA Training Momentum in Federated Fine-Tuning
von: Yan, Peishen, et al.
Veröffentlicht: (2026)
von: Yan, Peishen, et al.
Veröffentlicht: (2026)
VeriLeaky: Navigating IP Protection vs Utility in Fine-Tuning for LLM-Driven Verilog Coding
von: Wang, Zeng, et al.
Veröffentlicht: (2025)
von: Wang, Zeng, et al.
Veröffentlicht: (2025)
VAULT: Vigilant Adversarial Updates via LLM-Driven Retrieval-Augmented Generation for NLI
von: Kazoom, Roie, et al.
Veröffentlicht: (2025)
von: Kazoom, Roie, et al.
Veröffentlicht: (2025)
LISAA: A Framework for Large Language Model Information Security Awareness Assessment
von: Cohen, Ofir, et al.
Veröffentlicht: (2024)
von: Cohen, Ofir, et al.
Veröffentlicht: (2024)
Rotation-Preserving Supervised Fine-Tuning
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026)
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DOMBA: Double Model Balancing for Access-Controlled Language Models via Minimum-Bounded Aggregation
von: Segal, Tom, et al.
Veröffentlicht: (2024) -
LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
von: Schwartz, Yuval, et al.
Veröffentlicht: (2024) -
Addressing Key Challenges of Adversarial Attacks and Defenses in the Tabular Domain: A Methodological Framework for Coherence and Consistency
von: Itzhakev, Yael, et al.
Veröffentlicht: (2024) -
RAPID: Robust APT Detection and Investigation Using Context-Aware Deep Learning
von: Amaru, Yonatan, et al.
Veröffentlicht: (2024) -
Real-World Adversarial Attacks on RF-Based Drone Detectors
von: Gazit, Omer, et al.
Veröffentlicht: (2025)