The bitter lesson of misuse detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mariaccia, Hadrien, Segerie, Charbel-Raphaël, Dorn, Diego |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
von: Dorn, Diego, et al.
Veröffentlicht: (2024)
von: Dorn, Diego, et al.
Veröffentlicht: (2024)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
von: Wang, Liwen, et al.
Veröffentlicht: (2025)
von: Wang, Liwen, et al.
Veröffentlicht: (2025)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025)
Leaking LoRa: An Evaluation of Password Leaks and Knowledge Storage in Large Language Models
von: Marinelli, Ryan, et al.
Veröffentlicht: (2025)
von: Marinelli, Ryan, et al.
Veröffentlicht: (2025)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
Strategic Deflection: Defending LLMs from Logit Manipulation
von: Rachidy, Yassine, et al.
Veröffentlicht: (2025)
von: Rachidy, Yassine, et al.
Veröffentlicht: (2025)
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
von: Schnabl, Christoph, et al.
Veröffentlicht: (2025)
von: Schnabl, Christoph, et al.
Veröffentlicht: (2025)
Jailbreaking Large Language Models Through Content Concretization
von: Wahréus, Johan, et al.
Veröffentlicht: (2025)
von: Wahréus, Johan, et al.
Veröffentlicht: (2025)
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
von: Green, Tommaso, et al.
Veröffentlicht: (2025)
von: Green, Tommaso, et al.
Veröffentlicht: (2025)
Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation
von: Shao, Shuo, et al.
Veröffentlicht: (2025)
von: Shao, Shuo, et al.
Veröffentlicht: (2025)
A Survey on Agentic Security: Applications, Threats and Defenses
von: Shahriar, Asif, et al.
Veröffentlicht: (2025)
von: Shahriar, Asif, et al.
Veröffentlicht: (2025)
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
von: Liang, Zi, et al.
Veröffentlicht: (2025)
von: Liang, Zi, et al.
Veröffentlicht: (2025)
VelLMes: A high-interaction AI-based deception framework
von: Sladić, Muris, et al.
Veröffentlicht: (2025)
von: Sladić, Muris, et al.
Veröffentlicht: (2025)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
von: Liu, Yule, et al.
Veröffentlicht: (2025)
von: Liu, Yule, et al.
Veröffentlicht: (2025)
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
von: Yan, Yu, et al.
Veröffentlicht: (2025)
von: Yan, Yu, et al.
Veröffentlicht: (2025)
Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents
von: Liang, Zhibo, et al.
Veröffentlicht: (2025)
von: Liang, Zhibo, et al.
Veröffentlicht: (2025)
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
von: Fastowski, Alina, et al.
Veröffentlicht: (2025)
von: Fastowski, Alina, et al.
Veröffentlicht: (2025)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
QueryAttack: Jailbreaking Aligned Large Language Models Using Structured Non-natural Query Language
von: Zou, Qingsong, et al.
Veröffentlicht: (2025)
von: Zou, Qingsong, et al.
Veröffentlicht: (2025)
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
von: Leong, Chak Tou, et al.
Veröffentlicht: (2025)
von: Leong, Chak Tou, et al.
Veröffentlicht: (2025)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
MURMUR: Using cross-user chatter to break collaborative language agents in groups
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
von: Patlan, Atharv Singh, et al.
Veröffentlicht: (2025)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
von: Hu, Xuhao, et al.
Veröffentlicht: (2025)
von: Hu, Xuhao, et al.
Veröffentlicht: (2025)
LLM Jailbreak Detection for (Almost) Free!
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
von: Teja, Lekkala Sai, et al.
Veröffentlicht: (2025)
von: Teja, Lekkala Sai, et al.
Veröffentlicht: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
von: Mu, Junjie, et al.
Veröffentlicht: (2025)
von: Mu, Junjie, et al.
Veröffentlicht: (2025)
EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
von: Zhou, Xinyun, et al.
Veröffentlicht: (2025)
von: Zhou, Xinyun, et al.
Veröffentlicht: (2025)
PARASITE: Conditional System Prompt Poisoning to Hijack LLMs
von: Pham, Viet, et al.
Veröffentlicht: (2025)
von: Pham, Viet, et al.
Veröffentlicht: (2025)
Towards a scalable AI-driven framework for data-independent Cyber Threat Intelligence Information Extraction
von: Sorokoletova, Olga, et al.
Veröffentlicht: (2025)
von: Sorokoletova, Olga, et al.
Veröffentlicht: (2025)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
von: Ji, Wence, et al.
Veröffentlicht: (2025)
von: Ji, Wence, et al.
Veröffentlicht: (2025)
Defend LLMs Through Self-Consciousness
von: Huang, Boshi, et al.
Veröffentlicht: (2025)
von: Huang, Boshi, et al.
Veröffentlicht: (2025)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
CATMark: A Context-Aware Thresholding Framework for Robust Cross-Task Watermarking in Large Language Models
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
von: Dorn, Diego, et al.
Veröffentlicht: (2024) -
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
von: Zhao, Yi, et al.
Veröffentlicht: (2025) -
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
von: Wang, Liwen, et al.
Veröffentlicht: (2025) -
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
von: Akbar-Tajari, Mohammad, et al.
Veröffentlicht: (2025) -
Leaking LoRa: An Evaluation of Password Leaks and Knowledge Storage in Large Language Models
von: Marinelli, Ryan, et al.
Veröffentlicht: (2025)