Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dabas, Mahavir, Huynh, Tran, Billa, Nikhil Reddy, Wang, Jiachen T., Gao, Peng, Peris, Charith, Ma, Yao, Gupta, Rahul, Jin, Ming, Mittal, Prateek, Jia, Ruoxi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Memory-Induced Tool-Drift in LLM Agents
von: Dabas, Mahavir, et al.
Veröffentlicht: (2026)
von: Dabas, Mahavir, et al.
Veröffentlicht: (2026)
Déjà Vu
von: Houppermans, Sjef, et al.
Veröffentlicht: (2016)
von: Houppermans, Sjef, et al.
Veröffentlicht: (2016)
Efficient Data Shapley for Weighted Nearest Neighbor Algorithms
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024)
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024)
Characterizing Model-Native Skills
von: Kang, Feiyang, et al.
Veröffentlicht: (2026)
von: Kang, Feiyang, et al.
Veröffentlicht: (2026)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
von: Dabas, Mahavir, et al.
Veröffentlicht: (2025)
von: Dabas, Mahavir, et al.
Veröffentlicht: (2025)
Deja Vu from the Bridge.
von: Sullivan, Peggy
Veröffentlicht: (1979)
von: Sullivan, Peggy
Veröffentlicht: (1979)
DiPT: Enhancing LLM reasoning through diversified perspective-taking
von: Just, Hoang Anh, et al.
Veröffentlicht: (2024)
von: Just, Hoang Anh, et al.
Veröffentlicht: (2024)
Data Shapley in One Training Run
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024)
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024)
Déjà Vu Memorization in Vision-Language Models
von: Jayaraman, Bargav, et al.
Veröffentlicht: (2024)
von: Jayaraman, Bargav, et al.
Veröffentlicht: (2024)
The Electronic Revolution in Libraries: Microfilm Deja Vu?
von: Cady, Susan A.
Veröffentlicht: (1990)
von: Cady, Susan A.
Veröffentlicht: (1990)
Info Lit 2.0 or Déjà Vu?
von: Ianuzzi, Patricia Anne
Veröffentlicht: (2013)
von: Ianuzzi, Patricia Anne
Veröffentlicht: (2013)
Retracing the Past: LLMs Emit Training Data When They Get Lost
von: Ko, Myeongseob, et al.
Veröffentlicht: (2025)
von: Ko, Myeongseob, et al.
Veröffentlicht: (2025)
Capturing the Temporal Dependence of Training Data Influence
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024)
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024)
DejaVu: A Minimalistic Mechanism for Distributed Plurality Consensus
von: d'Amore, Francesco, et al.
Veröffentlicht: (2026)
von: d'Amore, Francesco, et al.
Veröffentlicht: (2026)
Déjà Vu? Decoding Repeated Reading from Eye Movements
von: Meiri, Yoav, et al.
Veröffentlicht: (2025)
von: Meiri, Yoav, et al.
Veröffentlicht: (2025)
News Deja Vu: Connecting Past and Present with Semantic Search
von: Franklin, Brevin, et al.
Veröffentlicht: (2024)
von: Franklin, Brevin, et al.
Veröffentlicht: (2024)
School and Public Library Relationships: Deja Vu or New Beginnings.
von: Fitzgibbons, Shirley A.
Veröffentlicht: (2001)
von: Fitzgibbons, Shirley A.
Veröffentlicht: (2001)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
von: Strati, Foteini, et al.
Veröffentlicht: (2024)
von: Strati, Foteini, et al.
Veröffentlicht: (2024)
Déjà Vu Packing: Optimizing FPGA Logic Clustering Runtime via Pattern Memoization
von: Liebster, Milo, et al.
Veröffentlicht: (2026)
von: Liebster, Milo, et al.
Veröffentlicht: (2026)
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
von: Wang, Jiachen T., et al.
Veröffentlicht: (2025)
von: Wang, Jiachen T., et al.
Veröffentlicht: (2025)
Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
von: Hwang, Jinwoo, et al.
Veröffentlicht: (2025)
von: Hwang, Jinwoo, et al.
Veröffentlicht: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
A Linguistic Analysis of Spontaneous Thoughts: Investigating Experiences of Déjà Vu, Unexpected Thoughts, and Involuntary Autobiographical Memories
von: Venkatesha, Videep, et al.
Veröffentlicht: (2025)
von: Venkatesha, Videep, et al.
Veröffentlicht: (2025)
Deja Vu in Plots: Leveraging Cross-Session Evidence with Retrieval-Augmented LLMs for Live Streaming Risk Assessment
von: Qiao, Yiran, et al.
Veröffentlicht: (2026)
von: Qiao, Yiran, et al.
Veröffentlicht: (2026)
BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
von: Liang, Shuang, et al.
Veröffentlicht: (2025)
von: Liang, Shuang, et al.
Veröffentlicht: (2025)
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
PatchDEMUX: A Certifiably Robust Framework for Multi-label Classifiers Against Adversarial Patches
von: Jacob, Dennis, et al.
Veröffentlicht: (2025)
von: Jacob, Dennis, et al.
Veröffentlicht: (2025)
The existence and controllability of nonautonomous system influenced by impulses on both state and control
von: Gupta, Garima, et al.
Veröffentlicht: (2024)
von: Gupta, Garima, et al.
Veröffentlicht: (2024)
Existence And Approximate Controllability for a class of Fractional Order Hemivariational Inequalities
von: Gupta, Garima, et al.
Veröffentlicht: (2024)
von: Gupta, Garima, et al.
Veröffentlicht: (2024)
Study on Control Problem of a Impulsive Neutral Integro-Differential Equations with Fading Memory
von: Gupta, Garima, et al.
Veröffentlicht: (2025)
von: Gupta, Garima, et al.
Veröffentlicht: (2025)
Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models
von: Zhao, Yunhan, et al.
Veröffentlicht: (2025)
von: Zhao, Yunhan, et al.
Veröffentlicht: (2025)
Position: Towards Resilience Against Adversarial Examples
von: Dai, Sihui, et al.
Veröffentlicht: (2024)
von: Dai, Sihui, et al.
Veröffentlicht: (2024)
K-Edit: Language Model Editing with Contextual Knowledge Awareness
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
von: Chen, Kejia, et al.
Veröffentlicht: (2026)
von: Chen, Kejia, et al.
Veröffentlicht: (2026)
Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual Attacks
von: Xie, Yong, et al.
Veröffentlicht: (2024)
von: Xie, Yong, et al.
Veröffentlicht: (2024)
Defenses Against Prompt Attacks Learn Surface Heuristics
von: Li, Shawn, et al.
Veröffentlicht: (2026)
von: Li, Shawn, et al.
Veröffentlicht: (2026)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
Déjà vu
von: Ernesto Raabe
Veröffentlicht: (2009)
von: Ernesto Raabe
Veröffentlicht: (2009)
Ähnliche Einträge
-
Memory-Induced Tool-Drift in LLM Agents
von: Dabas, Mahavir, et al.
Veröffentlicht: (2026) -
Déjà Vu
von: Houppermans, Sjef, et al.
Veröffentlicht: (2016) -
Efficient Data Shapley for Weighted Nearest Neighbor Algorithms
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024) -
Characterizing Model-Native Skills
von: Kang, Feiyang, et al.
Veröffentlicht: (2026) -
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
von: Dabas, Mahavir, et al.
Veröffentlicht: (2025)