Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kulkarni, Prashant, Namer, Assaf |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
von: Hossain, S M Asif, et al.
Veröffentlicht: (2025)
von: Hossain, S M Asif, et al.
Veröffentlicht: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
A Zero Trust Framework for Realization and Defense Against Generative AI Attacks in Power Grid
von: Munir, Md. Shirajum, et al.
Veröffentlicht: (2024)
von: Munir, Md. Shirajum, et al.
Veröffentlicht: (2024)
PECAN: A Deterministic Certified Defense Against Backdoor Attacks
von: Zhang, Yuhao, et al.
Veröffentlicht: (2023)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2023)
A No-Defense Defense Against Gradient-Based Adversarial Attacks on ML-NIDS: Is Less More?
von: elShehaby, Mohamed, et al.
Veröffentlicht: (2026)
von: elShehaby, Mohamed, et al.
Veröffentlicht: (2026)
A Systematic Review of Poisoning Attacks Against Large Language Models
von: Fendley, Neil, et al.
Veröffentlicht: (2025)
von: Fendley, Neil, et al.
Veröffentlicht: (2025)
A Defensive Framework Against Adversarial Attacks on Machine Learning-Based Network Intrusion Detection Systems
von: Tafreshian, Benyamin, et al.
Veröffentlicht: (2025)
von: Tafreshian, Benyamin, et al.
Veröffentlicht: (2025)
Dashed Line Defense: Plug-And-Play Defense Against Adaptive Score-Based Query Attacks
von: Fu, Yanzhang, et al.
Veröffentlicht: (2026)
von: Fu, Yanzhang, et al.
Veröffentlicht: (2026)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
Attacks and Defenses Against LLM Fingerprinting
von: Kurian, Kevin, et al.
Veröffentlicht: (2025)
von: Kurian, Kevin, et al.
Veröffentlicht: (2025)
SecureLearn -- An Attack-agnostic Defense for Multiclass Machine Learning Against Data Poisoning Attacks
von: Paracha, Anum, et al.
Veröffentlicht: (2025)
von: Paracha, Anum, et al.
Veröffentlicht: (2025)
A Survey on Model Extraction Attacks and Defenses for Large Language Models
von: Zhao, Kaixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Kaixiang, et al.
Veröffentlicht: (2025)
RAIFLE: Reconstruction Attacks on Interaction-based Federated Learning with Adversarial Data Manipulation
von: Pham, Dzung, et al.
Veröffentlicht: (2023)
von: Pham, Dzung, et al.
Veröffentlicht: (2023)
Defending Large Language Models Against Attacks With Residual Stream Activation Analysis
von: Kawasaki, Amelia, et al.
Veröffentlicht: (2024)
von: Kawasaki, Amelia, et al.
Veröffentlicht: (2024)
KDk: A Defense Mechanism Against Label Inference Attacks in Vertical Federated Learning
von: Arazzi, Marco, et al.
Veröffentlicht: (2024)
von: Arazzi, Marco, et al.
Veröffentlicht: (2024)
Optimal Defenses Against Gradient Reconstruction Attacks
von: Chen, Yuxiao, et al.
Veröffentlicht: (2024)
von: Chen, Yuxiao, et al.
Veröffentlicht: (2024)
MISLEAD: Manipulating Importance of Selected features for Learning Epsilon in Evasion Attack Deception
von: Khazanchi, Vidit, et al.
Veröffentlicht: (2024)
von: Khazanchi, Vidit, et al.
Veröffentlicht: (2024)
Membership Inference Attacks for Retrieval Based In-Context Learning for Document Question Answering
von: Kulkarni, Tejas, et al.
Veröffentlicht: (2026)
von: Kulkarni, Tejas, et al.
Veröffentlicht: (2026)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
Attack and Defense of Deep Learning Models in the Field of Web Attack Detection
von: Shi, Lijia, et al.
Veröffentlicht: (2024)
von: Shi, Lijia, et al.
Veröffentlicht: (2024)
Learning to Poison Large Language Models for Downstream Manipulation
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
von: Guo, Qiming, et al.
Veröffentlicht: (2025)
von: Guo, Qiming, et al.
Veröffentlicht: (2025)
Hijacking Large Language Models via Adversarial In-Context Learning
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)
Dummy-Aware Weighted Attack (DAWA): Breaking the Safe Sink in Dummy Class Defenses
von: Yu, Yunrui, et al.
Veröffentlicht: (2026)
von: Yu, Yunrui, et al.
Veröffentlicht: (2026)
A New Federated Learning Framework Against Gradient Inversion Attacks
von: Guo, Pengxin, et al.
Veröffentlicht: (2024)
von: Guo, Pengxin, et al.
Veröffentlicht: (2024)
A Taxonomy of Attacks and Defenses in Split Learning
von: Shabbir, Aqsa, et al.
Veröffentlicht: (2025)
von: Shabbir, Aqsa, et al.
Veröffentlicht: (2025)
MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models
von: He, Junhao, et al.
Veröffentlicht: (2025)
von: He, Junhao, et al.
Veröffentlicht: (2025)
ADAGE: Active Defenses Against GNN Extraction
von: Xu, Jing, et al.
Veröffentlicht: (2025)
von: Xu, Jing, et al.
Veröffentlicht: (2025)
CAPoW: Context-Aware AI-Assisted Proof of Work based DDoS Defense
von: Chakraborty, Trisha, et al.
Veröffentlicht: (2023)
von: Chakraborty, Trisha, et al.
Veröffentlicht: (2023)
FedMID: A Data-Free Method for Using Intermediate Outputs as a Defense Mechanism Against Poisoning Attacks in Federated Learning
von: Han, Sungwon, et al.
Veröffentlicht: (2024)
von: Han, Sungwon, et al.
Veröffentlicht: (2024)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
How to Defend Against Large-scale Model Poisoning Attacks in Federated Learning: A Vertical Solution
von: Wang, Jinbo, et al.
Veröffentlicht: (2024)
von: Wang, Jinbo, et al.
Veröffentlicht: (2024)
Data Reconstruction Attacks and Defenses: A Systematic Evaluation
von: Liu, Sheng, et al.
Veröffentlicht: (2024)
von: Liu, Sheng, et al.
Veröffentlicht: (2024)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
von: Panebianco, Francesco, et al.
Veröffentlicht: (2025)
von: Panebianco, Francesco, et al.
Veröffentlicht: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
Evaluating Apple Intelligence's Writing Tools for Privacy Against Large Language Model-Based Inference Attacks: Insights from Early Datasets
von: Soumik, Mohd. Farhan Israk, et al.
Veröffentlicht: (2025)
von: Soumik, Mohd. Farhan Israk, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
von: Hossain, S M Asif, et al.
Veröffentlicht: (2025) -
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024) -
A Zero Trust Framework for Realization and Defense Against Generative AI Attacks in Power Grid
von: Munir, Md. Shirajum, et al.
Veröffentlicht: (2024) -
PECAN: A Deterministic Certified Defense Against Backdoor Attacks
von: Zhang, Yuhao, et al.
Veröffentlicht: (2023) -
A No-Defense Defense Against Gradient-Based Adversarial Attacks on ML-NIDS: Is Less More?
von: elShehaby, Mohamed, et al.
Veröffentlicht: (2026)