Capability-Based Scaling Trends for LLM-Based Red-Teaming
Fuente:
arXiv
Salvato in:
| Autori principali: | Panfilov, Alexander, Kassianik, Paul, Andriushchenko, Maksym, Geiping, Jonas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
di: Terekhov, Mikhail, et al.
Pubblicazione: (2025)
di: Terekhov, Mikhail, et al.
Pubblicazione: (2025)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
di: Panfilov, Alexander, et al.
Pubblicazione: (2026)
di: Panfilov, Alexander, et al.
Pubblicazione: (2026)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
di: Boreiko, Valentyn, et al.
Pubblicazione: (2024)
di: Boreiko, Valentyn, et al.
Pubblicazione: (2024)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
di: Zymet, Jesse, et al.
Pubblicazione: (2026)
di: Zymet, Jesse, et al.
Pubblicazione: (2026)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
di: Rando, Javier, et al.
Pubblicazione: (2024)
di: Rando, Javier, et al.
Pubblicazione: (2024)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
di: Mehrotra, Anay, et al.
Pubblicazione: (2023)
di: Mehrotra, Anay, et al.
Pubblicazione: (2023)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
di: Sharma, Mrinank, et al.
Pubblicazione: (2025)
di: Sharma, Mrinank, et al.
Pubblicazione: (2025)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
di: Dong, Jianshuo, et al.
Pubblicazione: (2025)
di: Dong, Jianshuo, et al.
Pubblicazione: (2025)
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
di: Zhao, Andrew, et al.
Pubblicazione: (2024)
di: Zhao, Andrew, et al.
Pubblicazione: (2024)
There Are No Silly Questions: Evaluation of Offline LLM Capabilities from a Turkish Perspective
di: Yilmaz, Edibe, et al.
Pubblicazione: (2026)
di: Yilmaz, Edibe, et al.
Pubblicazione: (2026)
Policy-Invisible Violations in LLM-Based Agents
di: Wu, Jie, et al.
Pubblicazione: (2026)
di: Wu, Jie, et al.
Pubblicazione: (2026)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
di: Dobre, David, et al.
Pubblicazione: (2025)
di: Dobre, David, et al.
Pubblicazione: (2025)
Early Signs of Steganographic Capabilities in Frontier LLMs
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
di: Lee, Hyomin, et al.
Pubblicazione: (2026)
di: Lee, Hyomin, et al.
Pubblicazione: (2026)
Gandalf the Red: Adaptive Security for LLMs
di: Pfister, Niklas, et al.
Pubblicazione: (2025)
di: Pfister, Niklas, et al.
Pubblicazione: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
di: Halawi, Danny, et al.
Pubblicazione: (2024)
di: Halawi, Danny, et al.
Pubblicazione: (2024)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
di: Zhou, Kaiwen, et al.
Pubblicazione: (2025)
di: Zhou, Kaiwen, et al.
Pubblicazione: (2025)
Query-Based Adversarial Prompt Generation
di: Hayase, Jonathan, et al.
Pubblicazione: (2024)
di: Hayase, Jonathan, et al.
Pubblicazione: (2024)
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
di: Tong, Xin, et al.
Pubblicazione: (2025)
di: Tong, Xin, et al.
Pubblicazione: (2025)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
di: Sinha, Anusha, et al.
Pubblicazione: (2025)
di: Sinha, Anusha, et al.
Pubblicazione: (2025)
Scaling Trends in Language Model Robustness
di: Howe, Nikolaus, et al.
Pubblicazione: (2024)
di: Howe, Nikolaus, et al.
Pubblicazione: (2024)
MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
di: Kan, Chun Yan Ryan, et al.
Pubblicazione: (2026)
di: Kan, Chun Yan Ryan, et al.
Pubblicazione: (2026)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
di: Yang, Xiaoxue, et al.
Pubblicazione: (2025)
di: Yang, Xiaoxue, et al.
Pubblicazione: (2025)
Federated In-Context LLM Agent Learning
di: Wu, Panlong, et al.
Pubblicazione: (2024)
di: Wu, Panlong, et al.
Pubblicazione: (2024)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
di: Zhu, Sicheng, et al.
Pubblicazione: (2024)
di: Zhu, Sicheng, et al.
Pubblicazione: (2024)
Certifying LLM Safety against Adversarial Prompting
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
Extracting Memorized Training Data via Decomposition
di: Su, Ellen, et al.
Pubblicazione: (2024)
di: Su, Ellen, et al.
Pubblicazione: (2024)
UniTSyn: A Large-Scale Dataset Capable of Enhancing the Prowess of Large Language Models for Program Testing
di: He, Yifeng, et al.
Pubblicazione: (2024)
di: He, Yifeng, et al.
Pubblicazione: (2024)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2025)
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2025)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
di: Zhang, Jiawei, et al.
Pubblicazione: (2025)
di: Zhang, Jiawei, et al.
Pubblicazione: (2025)
KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection
di: Li, Yuexin, et al.
Pubblicazione: (2024)
di: Li, Yuexin, et al.
Pubblicazione: (2024)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
di: Liang, Buyun, et al.
Pubblicazione: (2025)
di: Liang, Buyun, et al.
Pubblicazione: (2025)
LLM Cyber Evaluations Don't Capture Real-World Risk
di: Lukošiūtė, Kamilė, et al.
Pubblicazione: (2025)
di: Lukošiūtė, Kamilė, et al.
Pubblicazione: (2025)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
di: Wu, Xiaoyu, et al.
Pubblicazione: (2025)
di: Wu, Xiaoyu, et al.
Pubblicazione: (2025)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
di: Spracklen, Joseph, et al.
Pubblicazione: (2026)
di: Spracklen, Joseph, et al.
Pubblicazione: (2026)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
di: Wu, Tong, et al.
Pubblicazione: (2024)
di: Wu, Tong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
di: Terekhov, Mikhail, et al.
Pubblicazione: (2025) -
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
di: Panfilov, Alexander, et al.
Pubblicazione: (2026) -
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
di: Boreiko, Valentyn, et al.
Pubblicazione: (2024) -
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
di: Panfilov, Alexander, et al.
Pubblicazione: (2025) -
Adaptive Instruction Composition for Automated LLM Red-Teaming
di: Zymet, Jesse, et al.
Pubblicazione: (2026)