Teach LLMs to Phish: Stealing Private Information from Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Panda, Ashwinee, Choquette-Choo, Christopher A., Zhang, Zhengming, Yang, Yaoqing, Mittal, Prateek |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Privacy Auditing of Large Language Models
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
von: Panda, Ashwinee, et al.
Veröffentlicht: (2022)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2022)
Private Fine-tuning of Large Language Models with Zeroth-order Optimization
von: Tang, Xinyu, et al.
Veröffentlicht: (2024)
von: Tang, Xinyu, et al.
Veröffentlicht: (2024)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
von: Liu, Ken Ziyu, et al.
Veröffentlicht: (2025)
von: Liu, Ken Ziyu, et al.
Veröffentlicht: (2025)
KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection
von: Li, Yuexin, et al.
Veröffentlicht: (2024)
von: Li, Yuexin, et al.
Veröffentlicht: (2024)
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
von: Qi, Xiangyu, et al.
Veröffentlicht: (2024)
von: Qi, Xiangyu, et al.
Veröffentlicht: (2024)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
Correlated Noise Provably Beats Independent Noise for Differentially Private Learning
von: Choquette-Choo, Christopher A., et al.
Veröffentlicht: (2023)
von: Choquette-Choo, Christopher A., et al.
Veröffentlicht: (2023)
Watermark Stealing in Large Language Models
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
Stealing User Prompts from Mixture of Experts
von: Yona, Itay, et al.
Veröffentlicht: (2024)
von: Yona, Itay, et al.
Veröffentlicht: (2024)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
von: Batra, Shourya, et al.
Veröffentlicht: (2025)
von: Batra, Shourya, et al.
Veröffentlicht: (2025)
Are aligned neural networks adversarially aligned?
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
von: Cao, Bochuan, et al.
Veröffentlicht: (2025)
von: Cao, Bochuan, et al.
Veröffentlicht: (2025)
User Inference Attacks on Large Language Models
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
Every Character Counts: From Vulnerability to Defense in Phishing Detection
von: Chiper, Maria, et al.
Veröffentlicht: (2025)
von: Chiper, Maria, et al.
Veröffentlicht: (2025)
Learning to Diagnose Privately: DP-Powered LLMs for Radiology Report Classification
von: Bhattacharjee, Payel, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Payel, et al.
Veröffentlicht: (2025)
LLMs unlock new paths to monetizing exploits
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
Phishing Detection in the Gen-AI Era: Quantized LLMs vs Classical Models
von: Thapa, Jikesh, et al.
Veröffentlicht: (2025)
von: Thapa, Jikesh, et al.
Veröffentlicht: (2025)
MURMUR: Using cross-user chatter to break collaborative language agents in groups
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
MAPLE: Metadata Augmented Private Language Evolution
von: Chien, Eli, et al.
Veröffentlicht: (2026)
von: Chien, Eli, et al.
Veröffentlicht: (2026)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
von: Frikha, Ahmed, et al.
Veröffentlicht: (2024)
von: Frikha, Ahmed, et al.
Veröffentlicht: (2024)
SentinelLMs: Encrypted Input Adaptation and Fine-tuning of Language Models for Private and Secure Inference
von: Mishra, Abhijit, et al.
Veröffentlicht: (2023)
von: Mishra, Abhijit, et al.
Veröffentlicht: (2023)
Large Language Model Watermark Stealing With Mixed Integer Programming
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2024)
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2024)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
von: Lin, Shi, et al.
Veröffentlicht: (2024)
von: Lin, Shi, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models for Phishing Detection, Self-Consistency, Faithfulness, and Explainability
von: Kuikel, Shova, et al.
Veröffentlicht: (2025)
von: Kuikel, Shova, et al.
Veröffentlicht: (2025)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
An Explainable Transformer-based Model for Phishing Email Detection: A Large Language Model Approach
von: Uddin, Mohammad Amaz, et al.
Veröffentlicht: (2024)
von: Uddin, Mohammad Amaz, et al.
Veröffentlicht: (2024)
Personal Information Parroting in Language Models
von: Subramani, Nishant, et al.
Veröffentlicht: (2026)
von: Subramani, Nishant, et al.
Veröffentlicht: (2026)
Differentially Private Learning Needs Better Model Initialization and Self-Distillation
von: Ngong, Ivoline C., et al.
Veröffentlicht: (2024)
von: Ngong, Ivoline C., et al.
Veröffentlicht: (2024)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
von: Küchler, Nicolas, et al.
Veröffentlicht: (2025)
von: Küchler, Nicolas, et al.
Veröffentlicht: (2025)
NoPhish: Efficient Chrome Extension for Phishing Detection Using Machine Learning Techniques
von: Thaqi, Leand, et al.
Veröffentlicht: (2024)
von: Thaqi, Leand, et al.
Veröffentlicht: (2024)
Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
von: Borkar, Jaydeep, et al.
Veröffentlicht: (2025)
von: Borkar, Jaydeep, et al.
Veröffentlicht: (2025)
Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private
von: Wu, Ruihan, et al.
Veröffentlicht: (2025)
von: Wu, Ruihan, et al.
Veröffentlicht: (2025)
Exploring Query Efficient Data Generation towards Data-free Model Stealing in Hard Label Setting
von: Pei, Gaozheng, et al.
Veröffentlicht: (2024)
von: Pei, Gaozheng, et al.
Veröffentlicht: (2024)
VaultGemma: A Differentially Private Gemma Model
von: Sinha, Amer, et al.
Veröffentlicht: (2025)
von: Sinha, Amer, et al.
Veröffentlicht: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
PhreshPhish: A Real-World, High-Quality, Large-Scale Phishing Website Dataset and Benchmark
von: Dalton, Thomas, et al.
Veröffentlicht: (2025)
von: Dalton, Thomas, et al.
Veröffentlicht: (2025)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
Beyond A Fixed Seal: Adaptive Stealing Watermark in Large Language Models
von: Zhang, Shuhao, et al.
Veröffentlicht: (2026)
von: Zhang, Shuhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Privacy Auditing of Large Language Models
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025) -
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
von: Panda, Ashwinee, et al.
Veröffentlicht: (2022) -
Private Fine-tuning of Large Language Models with Zeroth-order Optimization
von: Tang, Xinyu, et al.
Veröffentlicht: (2024) -
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
von: Liu, Ken Ziyu, et al.
Veröffentlicht: (2025) -
KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection
von: Li, Yuexin, et al.
Veröffentlicht: (2024)