A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dobre, David, Mofakhami, Mehrnaz, Xhonneux, Sophie, Schwinn, Leo, Gidel, Gauthier |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-Safety Evaluations Lack Robustness
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
Efficient Adversarial Training in LLMs with Continuous Attacks
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024)
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024)
In-Context Learning Can Re-learn Forbidden Tasks
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024)
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
von: Schwinn, Leo, et al.
Veröffentlicht: (2026)
von: Schwinn, Leo, et al.
Veröffentlicht: (2026)
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
von: Schwinn, Leo, et al.
Veröffentlicht: (2024)
von: Schwinn, Leo, et al.
Veröffentlicht: (2024)
Adaptive Instruction Composition for Automated LLM Red-Teaming
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
von: Li, Xiaochen, et al.
Veröffentlicht: (2024)
von: Li, Xiaochen, et al.
Veröffentlicht: (2024)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
von: Lee, Seanie, et al.
Veröffentlicht: (2024)
von: Lee, Seanie, et al.
Veröffentlicht: (2024)
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
von: Rahman, Md Hafizur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Hafizur, et al.
Veröffentlicht: (2026)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
von: Liu, Kangwei, et al.
Veröffentlicht: (2025)
von: Liu, Kangwei, et al.
Veröffentlicht: (2025)
Gandalf the Red: Adaptive Security for LLMs
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting
von: Zhang, Hanxiu, et al.
Veröffentlicht: (2025)
von: Zhang, Hanxiu, et al.
Veröffentlicht: (2025)
Tight Lower Bounds and Improved Convergence in Performative Prediction
von: Khorsandi, Pedram, et al.
Veröffentlicht: (2024)
von: Khorsandi, Pedram, et al.
Veröffentlicht: (2024)
Interpreting the Repeated Token Phenomenon in Large Language Models
von: Yona, Itay, et al.
Veröffentlicht: (2025)
von: Yona, Itay, et al.
Veröffentlicht: (2025)
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
von: Ahmed, Mohamed, et al.
Veröffentlicht: (2025)
von: Ahmed, Mohamed, et al.
Veröffentlicht: (2025)
Fast Proxies for LLM Robustness Evaluation
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis
von: Mostafa, Ahmed, et al.
Veröffentlicht: (2025)
von: Mostafa, Ahmed, et al.
Veröffentlicht: (2025)
LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
von: Moia, Vitor Hugo Galhardo, et al.
Veröffentlicht: (2025)
von: Moia, Vitor Hugo Galhardo, et al.
Veröffentlicht: (2025)
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
von: Sharma, Mrinank, et al.
Veröffentlicht: (2025)
von: Sharma, Mrinank, et al.
Veröffentlicht: (2025)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
von: Kaunismaa, Jackson, et al.
Veröffentlicht: (2026)
von: Kaunismaa, Jackson, et al.
Veröffentlicht: (2026)
A Fast, Reliable, and Secure Programming Language for LLM Agents with Code Actions
von: Mell, Stephen, et al.
Veröffentlicht: (2025)
von: Mell, Stephen, et al.
Veröffentlicht: (2025)
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
von: Alhazbi, Saeif, et al.
Veröffentlicht: (2025)
von: Alhazbi, Saeif, et al.
Veröffentlicht: (2025)
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
von: Aldahoul, Nouar, et al.
Veröffentlicht: (2025)
von: Aldahoul, Nouar, et al.
Veröffentlicht: (2025)
Linking Cryptoasset Attribution Tags to Knowledge Graph Entities: An LLM-based Approach
von: Avice, Régnier, et al.
Veröffentlicht: (2025)
von: Avice, Régnier, et al.
Veröffentlicht: (2025)
PromptScreen: Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
von: Rao, Akshaj Prashanth, et al.
Veröffentlicht: (2025)
von: Rao, Akshaj Prashanth, et al.
Veröffentlicht: (2025)
Federated In-Context LLM Agent Learning
von: Wu, Panlong, et al.
Veröffentlicht: (2024)
von: Wu, Panlong, et al.
Veröffentlicht: (2024)
Rethinking How to Evaluate Language Model Jailbreak
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
von: Salahuddin, Salahuddin, et al.
Veröffentlicht: (2025)
von: Salahuddin, Salahuddin, et al.
Veröffentlicht: (2025)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
Policy-Invisible Violations in LLM-Based Agents
von: Wu, Jie, et al.
Veröffentlicht: (2026)
von: Wu, Jie, et al.
Veröffentlicht: (2026)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
von: Chrabąszcz, Maciej, et al.
Veröffentlicht: (2025)
von: Chrabąszcz, Maciej, et al.
Veröffentlicht: (2025)
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLM-Safety Evaluations Lack Robustness
von: Beyer, Tim, et al.
Veröffentlicht: (2025) -
Efficient Adversarial Training in LLMs with Continuous Attacks
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024) -
In-Context Learning Can Re-learn Forbidden Tasks
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024) -
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
von: Schwinn, Leo, et al.
Veröffentlicht: (2026) -
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
von: Schwinn, Leo, et al.
Veröffentlicht: (2024)