Coercing LLMs to do and reveal (almost) anything
Fuente:
arXiv
Salvato in:
| Autori principali: | Geiping, Jonas, Stein, Alex, Shu, Manli, Saifullah, Khalid, Wen, Yuxin, Goldstein, Tom |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On the Reliability of Watermarks for Large Language Models
di: Kirchenbauer, John, et al.
Pubblicazione: (2023)
di: Kirchenbauer, John, et al.
Pubblicazione: (2023)
A Watermark for Large Language Models
di: Kirchenbauer, John, et al.
Pubblicazione: (2023)
di: Kirchenbauer, John, et al.
Pubblicazione: (2023)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
di: Wen, Yuxin, et al.
Pubblicazione: (2024)
di: Wen, Yuxin, et al.
Pubblicazione: (2024)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
di: Boreiko, Valentyn, et al.
Pubblicazione: (2024)
di: Boreiko, Valentyn, et al.
Pubblicazione: (2024)
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
di: Wallace, Eric, et al.
Pubblicazione: (2024)
di: Wallace, Eric, et al.
Pubblicazione: (2024)
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
di: Xiao, Yuxin, et al.
Pubblicazione: (2023)
di: Xiao, Yuxin, et al.
Pubblicazione: (2023)
How Vulnerable Are Edge LLMs?
di: Ding, Ao, et al.
Pubblicazione: (2026)
di: Ding, Ao, et al.
Pubblicazione: (2026)
UCD: Unlearning in LLMs via Contrastive Decoding
di: Suriyakumar, Vinith M., et al.
Pubblicazione: (2025)
di: Suriyakumar, Vinith M., et al.
Pubblicazione: (2025)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
di: Dotsinski, Asen, et al.
Pubblicazione: (2026)
di: Dotsinski, Asen, et al.
Pubblicazione: (2026)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
di: Zhao, Xuandong, et al.
Pubblicazione: (2024)
di: Zhao, Xuandong, et al.
Pubblicazione: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
di: Sivapiromrat, Sanhanat, et al.
Pubblicazione: (2025)
di: Sivapiromrat, Sanhanat, et al.
Pubblicazione: (2025)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
di: Hasan, Adib, et al.
Pubblicazione: (2024)
di: Hasan, Adib, et al.
Pubblicazione: (2024)
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
di: Mathew, Yohan, et al.
Pubblicazione: (2024)
di: Mathew, Yohan, et al.
Pubblicazione: (2024)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
di: Brown, Hannah, et al.
Pubblicazione: (2024)
di: Brown, Hannah, et al.
Pubblicazione: (2024)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
di: Price, Sara, et al.
Pubblicazione: (2024)
di: Price, Sara, et al.
Pubblicazione: (2024)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
di: Wang, Linlin, et al.
Pubblicazione: (2025)
di: Wang, Linlin, et al.
Pubblicazione: (2025)
HARMONIC: Harnessing LLMs for Tabular Data Synthesis and Privacy Protection
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
LLM Unlearning Should Be Form-Independent
di: Ye, Xiaotian, et al.
Pubblicazione: (2025)
di: Ye, Xiaotian, et al.
Pubblicazione: (2025)
Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs
di: Joshi, Kunj, et al.
Pubblicazione: (2025)
di: Joshi, Kunj, et al.
Pubblicazione: (2025)
Can We Infer Confidential Properties of Training Data from LLMs?
di: Huang, Pengrun, et al.
Pubblicazione: (2025)
di: Huang, Pengrun, et al.
Pubblicazione: (2025)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
di: Suri, Omar Farooq Khan, et al.
Pubblicazione: (2025)
di: Suri, Omar Farooq Khan, et al.
Pubblicazione: (2025)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
di: Li, Ran, et al.
Pubblicazione: (2025)
di: Li, Ran, et al.
Pubblicazione: (2025)
Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion
di: Souri, Hossein, et al.
Pubblicazione: (2024)
di: Souri, Hossein, et al.
Pubblicazione: (2024)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
di: Roth, Tom, et al.
Pubblicazione: (2021)
di: Roth, Tom, et al.
Pubblicazione: (2021)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
di: Meeus, Matthieu, et al.
Pubblicazione: (2024)
di: Meeus, Matthieu, et al.
Pubblicazione: (2024)
Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs
di: Atwal, Tevin, et al.
Pubblicazione: (2025)
di: Atwal, Tevin, et al.
Pubblicazione: (2025)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
di: Xu, Yuancheng, et al.
Pubblicazione: (2024)
di: Xu, Yuancheng, et al.
Pubblicazione: (2024)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
di: Liu, Bing, et al.
Pubblicazione: (2026)
di: Liu, Bing, et al.
Pubblicazione: (2026)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
di: Yoon, Do-hyeon, et al.
Pubblicazione: (2025)
di: Yoon, Do-hyeon, et al.
Pubblicazione: (2025)
LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
di: Schwartz, Yuval, et al.
Pubblicazione: (2024)
di: Schwartz, Yuval, et al.
Pubblicazione: (2024)
Shh, don't say that! Domain Certification in LLMs
di: Emde, Cornelius, et al.
Pubblicazione: (2025)
di: Emde, Cornelius, et al.
Pubblicazione: (2025)
LMO-DP: Optimizing the Randomization Mechanism for Differentially Private Fine-Tuning (Large) Language Models
di: Yang, Qin, et al.
Pubblicazione: (2024)
di: Yang, Qin, et al.
Pubblicazione: (2024)
Private prediction for large-scale synthetic text generation
di: Amin, Kareem, et al.
Pubblicazione: (2024)
di: Amin, Kareem, et al.
Pubblicazione: (2024)
Training Data Reconstruction: Privacy due to Uncertainty?
di: Runkel, Christina, et al.
Pubblicazione: (2024)
di: Runkel, Christina, et al.
Pubblicazione: (2024)
ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?
di: Liu, Peihan, et al.
Pubblicazione: (2026)
di: Liu, Peihan, et al.
Pubblicazione: (2026)
TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection
di: Sander, Tom, et al.
Pubblicazione: (2026)
di: Sander, Tom, et al.
Pubblicazione: (2026)
Documenti analoghi
-
On the Reliability of Watermarks for Large Language Models
di: Kirchenbauer, John, et al.
Pubblicazione: (2023) -
A Watermark for Large Language Models
di: Kirchenbauer, John, et al.
Pubblicazione: (2023) -
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
di: Wen, Yuxin, et al.
Pubblicazione: (2024) -
Capability-Based Scaling Trends for LLM-Based Red-Teaming
di: Panfilov, Alexander, et al.
Pubblicazione: (2025) -
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
di: Boreiko, Valentyn, et al.
Pubblicazione: (2024)