LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
Fuente:
arXiv
Salvato in:
| Autori principali: | Hilel, Almog, Bhagwat, Riddhi, Shenfeld, Idan, Andreas, Jacob, Choshen, Leshem |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
User Inference Attacks on Large Language Models
di: Kandpal, Nikhil, et al.
Pubblicazione: (2023)
di: Kandpal, Nikhil, et al.
Pubblicazione: (2023)
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
di: Chen, Bocheng, et al.
Pubblicazione: (2024)
di: Chen, Bocheng, et al.
Pubblicazione: (2024)
Mind the Privacy Unit! User-Level Differential Privacy for Language Model Fine-Tuning
di: Chua, Lynn, et al.
Pubblicazione: (2024)
di: Chua, Lynn, et al.
Pubblicazione: (2024)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
di: Damani, Mehul, et al.
Pubblicazione: (2025)
di: Damani, Mehul, et al.
Pubblicazione: (2025)
Stealing User Prompts from Mixture of Experts
di: Yona, Itay, et al.
Pubblicazione: (2024)
di: Yona, Itay, et al.
Pubblicazione: (2024)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
di: Hines, Keegan, et al.
Pubblicazione: (2024)
di: Hines, Keegan, et al.
Pubblicazione: (2024)
An Early Categorization of Prompt Injection Attacks on Large Language Models
di: Rossi, Sippo, et al.
Pubblicazione: (2024)
di: Rossi, Sippo, et al.
Pubblicazione: (2024)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
di: Suri, Omar Farooq Khan, et al.
Pubblicazione: (2025)
di: Suri, Omar Farooq Khan, et al.
Pubblicazione: (2025)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
di: Yan, Jun, et al.
Pubblicazione: (2023)
di: Yan, Jun, et al.
Pubblicazione: (2023)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
di: Yao, Duanyi, et al.
Pubblicazione: (2026)
di: Yao, Duanyi, et al.
Pubblicazione: (2026)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
di: Ahmed, Mohamed, et al.
Pubblicazione: (2025)
di: Ahmed, Mohamed, et al.
Pubblicazione: (2025)
Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
di: Kimura, Subaru, et al.
Pubblicazione: (2024)
di: Kimura, Subaru, et al.
Pubblicazione: (2024)
User-Side Realization
di: Sato, Ryoma
Pubblicazione: (2024)
di: Sato, Ryoma
Pubblicazione: (2024)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
di: Jia, Feiran, et al.
Pubblicazione: (2024)
di: Jia, Feiran, et al.
Pubblicazione: (2024)
Mitigating Data Injection Attacks on Federated Learning
di: Shalom, Or, et al.
Pubblicazione: (2023)
di: Shalom, Or, et al.
Pubblicazione: (2023)
Textual analysis of End User License Agreement for red-flagging potentially malicious software
di: Khan, Behraj, et al.
Pubblicazione: (2024)
di: Khan, Behraj, et al.
Pubblicazione: (2024)
DistilLock: Safeguarding LLMs from Unauthorized Knowledge Distillation on the Edge
di: Mohanty, Asmita, et al.
Pubblicazione: (2025)
di: Mohanty, Asmita, et al.
Pubblicazione: (2025)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
di: Correia, Pedro H. Barcha, et al.
Pubblicazione: (2026)
di: Correia, Pedro H. Barcha, et al.
Pubblicazione: (2026)
Reformulation is All You Need: Addressing Malicious Text Features in DNNs
di: Jiang, Yi, et al.
Pubblicazione: (2025)
di: Jiang, Yi, et al.
Pubblicazione: (2025)
Can Gradient Descent Simulate Prompting?
di: Zhang, Eric, et al.
Pubblicazione: (2025)
di: Zhang, Eric, et al.
Pubblicazione: (2025)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
di: Yoon, Do-hyeon, et al.
Pubblicazione: (2025)
di: Yoon, Do-hyeon, et al.
Pubblicazione: (2025)
LLM Unlearning Should Be Form-Independent
di: Ye, Xiaotian, et al.
Pubblicazione: (2025)
di: Ye, Xiaotian, et al.
Pubblicazione: (2025)
GCG Attack On A Diffusion LLM
di: Neyroud, Ruben, et al.
Pubblicazione: (2025)
di: Neyroud, Ruben, et al.
Pubblicazione: (2025)
Differentially Private Knowledge Distillation via Synthetic Text Generation
di: Flemings, James, et al.
Pubblicazione: (2024)
di: Flemings, James, et al.
Pubblicazione: (2024)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
di: Wang, Linlin, et al.
Pubblicazione: (2025)
di: Wang, Linlin, et al.
Pubblicazione: (2025)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
di: Cui, Xinyue, et al.
Pubblicazione: (2025)
di: Cui, Xinyue, et al.
Pubblicazione: (2025)
Fine-Tuning Large Language Models with User-Level Differential Privacy
di: Charles, Zachary, et al.
Pubblicazione: (2024)
di: Charles, Zachary, et al.
Pubblicazione: (2024)
LLMGuard: Guarding Against Unsafe LLM Behavior
di: Goyal, Shubh, et al.
Pubblicazione: (2024)
di: Goyal, Shubh, et al.
Pubblicazione: (2024)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
di: Assogba, Yannick, et al.
Pubblicazione: (2026)
di: Assogba, Yannick, et al.
Pubblicazione: (2026)
Localizing Malicious Outputs from CodeLLM
di: Borana, Mayukh, et al.
Pubblicazione: (2025)
di: Borana, Mayukh, et al.
Pubblicazione: (2025)
Architecture Matters: Comparing RAG Systems under Knowledge Base Poisoning
di: Korn, Samuel
Pubblicazione: (2026)
di: Korn, Samuel
Pubblicazione: (2026)
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
di: Biswas, Sajib, et al.
Pubblicazione: (2025)
di: Biswas, Sajib, et al.
Pubblicazione: (2025)
PIArena: A Platform for Prompt Injection Evaluation
di: Geng, Runpeng, et al.
Pubblicazione: (2026)
di: Geng, Runpeng, et al.
Pubblicazione: (2026)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
di: Liu, Yupei, et al.
Pubblicazione: (2023)
di: Liu, Yupei, et al.
Pubblicazione: (2023)
Improving LLM Safety Alignment with Dual-Objective Optimization
di: Zhao, Xuandong, et al.
Pubblicazione: (2025)
di: Zhao, Xuandong, et al.
Pubblicazione: (2025)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
di: Liu, Ruixuan, et al.
Pubblicazione: (2026)
di: Liu, Ruixuan, et al.
Pubblicazione: (2026)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
di: Wang, Tianchun, et al.
Pubblicazione: (2024)
di: Wang, Tianchun, et al.
Pubblicazione: (2024)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
di: Krishna, Arjun, et al.
Pubblicazione: (2025)
di: Krishna, Arjun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
User Inference Attacks on Large Language Models
di: Kandpal, Nikhil, et al.
Pubblicazione: (2023) -
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
di: Chen, Bocheng, et al.
Pubblicazione: (2024) -
Mind the Privacy Unit! User-Level Differential Privacy for Language Model Fine-Tuning
di: Chua, Lynn, et al.
Pubblicazione: (2024) -
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
di: Damani, Mehul, et al.
Pubblicazione: (2025) -
Stealing User Prompts from Mixture of Experts
di: Yona, Itay, et al.
Pubblicazione: (2024)