Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Fastowski, Alina, Prenkaj, Bardh, Li, Yuxiao, Kasneci, Gjergji |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Confidence to Collapse in LLM Factual Robustness
di: Fastowski, Alina, et al.
Pubblicazione: (2025)
di: Fastowski, Alina, et al.
Pubblicazione: (2025)
Analysing the Safety Pitfalls of Steering Vectors
di: Li, Yuxiao, et al.
Pubblicazione: (2026)
di: Li, Yuxiao, et al.
Pubblicazione: (2026)
Understanding Knowledge Drift in LLMs through Misinformation
di: Fastowski, Alina, et al.
Pubblicazione: (2024)
di: Fastowski, Alina, et al.
Pubblicazione: (2024)
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
di: Kocak, Aysenur, et al.
Pubblicazione: (2025)
di: Kocak, Aysenur, et al.
Pubblicazione: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
di: Liu, Fan, et al.
Pubblicazione: (2024)
di: Liu, Fan, et al.
Pubblicazione: (2024)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
di: Yuan, Chenchen, et al.
Pubblicazione: (2026)
di: Yuan, Chenchen, et al.
Pubblicazione: (2026)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
di: Leemann, Tobias, et al.
Pubblicazione: (2024)
di: Leemann, Tobias, et al.
Pubblicazione: (2024)
Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models
di: Yuan, Chenchen, et al.
Pubblicazione: (2025)
di: Yuan, Chenchen, et al.
Pubblicazione: (2025)
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
di: Berezin, Sergey, et al.
Pubblicazione: (2025)
di: Berezin, Sergey, et al.
Pubblicazione: (2025)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
di: Baumgärtner, Tim, et al.
Pubblicazione: (2024)
di: Baumgärtner, Tim, et al.
Pubblicazione: (2024)
Emoti-Attack: Zero-Perturbation Adversarial Attacks on NLP Systems via Emoji Sequences
di: Zhang, Yangshijie
Pubblicazione: (2025)
di: Zhang, Yangshijie
Pubblicazione: (2025)
Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models
di: Liu, Xiao, et al.
Pubblicazione: (2024)
di: Liu, Xiao, et al.
Pubblicazione: (2024)
Adversarial Attacks Against Automated Fact-Checking: A Survey
di: Liu, Fanzhen, et al.
Pubblicazione: (2025)
di: Liu, Fanzhen, et al.
Pubblicazione: (2025)
Fast Adversarial Attacks on Language Models In One GPU Minute
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2024)
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2024)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
di: Pu, Rui, et al.
Pubblicazione: (2024)
di: Pu, Rui, et al.
Pubblicazione: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
di: Xu, Zhao, et al.
Pubblicazione: (2024)
di: Xu, Zhao, et al.
Pubblicazione: (2024)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
di: Teja, Lekkala Sai, et al.
Pubblicazione: (2025)
di: Teja, Lekkala Sai, et al.
Pubblicazione: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
di: Mu, Junjie, et al.
Pubblicazione: (2025)
di: Mu, Junjie, et al.
Pubblicazione: (2025)
LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
di: Li, Xingyu, et al.
Pubblicazione: (2025)
di: Li, Xingyu, et al.
Pubblicazione: (2025)
Hacc-Man: An Arcade Game for Jailbreaking LLMs
di: Valentim, Matheus, et al.
Pubblicazione: (2024)
di: Valentim, Matheus, et al.
Pubblicazione: (2024)
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
di: Pisano, Matthew, et al.
Pubblicazione: (2023)
di: Pisano, Matthew, et al.
Pubblicazione: (2023)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
di: Akbar-Tajari, Mohammad, et al.
Pubblicazione: (2025)
di: Akbar-Tajari, Mohammad, et al.
Pubblicazione: (2025)
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
di: Upadhayay, Bibek, et al.
Pubblicazione: (2024)
di: Upadhayay, Bibek, et al.
Pubblicazione: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
di: Ni, Ziyi, et al.
Pubblicazione: (2025)
di: Ni, Ziyi, et al.
Pubblicazione: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
di: Gohil, Vasudev
Pubblicazione: (2025)
di: Gohil, Vasudev
Pubblicazione: (2025)
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
di: Kirch, Nathalie, et al.
Pubblicazione: (2024)
di: Kirch, Nathalie, et al.
Pubblicazione: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
di: Wei, Jiali, et al.
Pubblicazione: (2026)
di: Wei, Jiali, et al.
Pubblicazione: (2026)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
di: Aldahoul, Nouar, et al.
Pubblicazione: (2025)
di: Aldahoul, Nouar, et al.
Pubblicazione: (2025)
The Resurgence of GCG Adversarial Attacks on Large Language Models
di: Tan, Yuting, et al.
Pubblicazione: (2025)
di: Tan, Yuting, et al.
Pubblicazione: (2025)
FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks
di: Xu, Naen, et al.
Pubblicazione: (2026)
di: Xu, Naen, et al.
Pubblicazione: (2026)
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
di: Yan, Dong, et al.
Pubblicazione: (2026)
di: Yan, Dong, et al.
Pubblicazione: (2026)
Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
di: Liang, Zi, et al.
Pubblicazione: (2025)
di: Liang, Zi, et al.
Pubblicazione: (2025)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
di: Lv, Lijia, et al.
Pubblicazione: (2024)
di: Lv, Lijia, et al.
Pubblicazione: (2024)
HauntAttack: When Attack Follows Reasoning as a Shadow
di: Ma, Jingyuan, et al.
Pubblicazione: (2025)
di: Ma, Jingyuan, et al.
Pubblicazione: (2025)
RAZOR: Sharpening Knowledge by Cutting Bias with Unsupervised Text Rewriting
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
di: Liu, Zesen, et al.
Pubblicazione: (2024)
di: Liu, Zesen, et al.
Pubblicazione: (2024)
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation
di: Shahariar, G M, et al.
Pubblicazione: (2024)
di: Shahariar, G M, et al.
Pubblicazione: (2024)
Documenti analoghi
-
From Confidence to Collapse in LLM Factual Robustness
di: Fastowski, Alina, et al.
Pubblicazione: (2025) -
Analysing the Safety Pitfalls of Steering Vectors
di: Li, Yuxiao, et al.
Pubblicazione: (2026) -
Understanding Knowledge Drift in LLMs through Misinformation
di: Fastowski, Alina, et al.
Pubblicazione: (2024) -
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
di: Kocak, Aysenur, et al.
Pubblicazione: (2025) -
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
di: Liu, Fan, et al.
Pubblicazione: (2024)