Scaling Patterns in Adversarial Alignment: Evidence from Multi-LLM Jailbreak Experiments
Fuente:
arXiv
Saved in:
| Main Authors: | Nathanson, Samuel, Williams, Rebecca, Matuszek, Cynthia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TRIZ Agents: A Multi-Agent LLM Approach for TRIZ-Based Innovation
by: Szczepanik, Kamil, et al.
Published: (2025)
by: Szczepanik, Kamil, et al.
Published: (2025)
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
by: Zhang, Peiyan, et al.
Published: (2025)
by: Zhang, Peiyan, et al.
Published: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
by: Pan, Leyi, et al.
Published: (2024)
by: Pan, Leyi, et al.
Published: (2024)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
A Semantic Invariant Robust Watermark for Large Language Models
by: Liu, Aiwei, et al.
Published: (2023)
by: Liu, Aiwei, et al.
Published: (2023)
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
by: Liu, Aiwei, et al.
Published: (2024)
by: Liu, Aiwei, et al.
Published: (2024)
GEML: A Grammar-based Evolutionary Machine Learning Approach for Design-Pattern Detection
by: Barbudo, Rafael, et al.
Published: (2024)
by: Barbudo, Rafael, et al.
Published: (2024)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
by: Radosky, Lukas, et al.
Published: (2026)
by: Radosky, Lukas, et al.
Published: (2026)
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
by: Chen, Junting, et al.
Published: (2024)
by: Chen, Junting, et al.
Published: (2024)
Aligning LLMs for Multilingual Consistency in Enterprise Applications
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
BudgetMLAgent: A Cost-Effective LLM Multi-Agent system for Automating Machine Learning Tasks
by: Gandhi, Shubham, et al.
Published: (2024)
by: Gandhi, Shubham, et al.
Published: (2024)
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
CAPE: Corrective Actions from Precondition Errors using Large Language Models
by: Raman, Shreyas Sundara, et al.
Published: (2022)
by: Raman, Shreyas Sundara, et al.
Published: (2022)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
by: Wang, Yanshu, et al.
Published: (2026)
by: Wang, Yanshu, et al.
Published: (2026)
PLUGH: A Benchmark for Spatial Understanding and Reasoning in Large Language Models
by: Tikhonov, Alexey
Published: (2024)
by: Tikhonov, Alexey
Published: (2024)
Pioneer Agent: Continual Improvement of Small Language Models in Production
by: Atreja, Dhruv, et al.
Published: (2026)
by: Atreja, Dhruv, et al.
Published: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
by: Wang, Xinhai, et al.
Published: (2026)
by: Wang, Xinhai, et al.
Published: (2026)
BreakFun: Jailbreaking LLMs via Schema Exploitation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
Collaborative LLM Agents for C4 Software Architecture Design Automation
by: Szczepanik, Kamil, et al.
Published: (2025)
by: Szczepanik, Kamil, et al.
Published: (2025)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
by: Kaiser, Daniel, et al.
Published: (2025)
by: Kaiser, Daniel, et al.
Published: (2025)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
Can AI Assist in Olympiad Coding
by: Ren, Samuel
Published: (2025)
by: Ren, Samuel
Published: (2025)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
by: Demir, M. Mikail, et al.
Published: (2025)
by: Demir, M. Mikail, et al.
Published: (2025)
Method for Aggregating Unstructured Data Using Large Language Models
by: Lazebnyi, Vsevolod, et al.
Published: (2026)
by: Lazebnyi, Vsevolod, et al.
Published: (2026)
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
PhenoAuth: A Novel PUF-Phenotype-based Authentication Protocol for IoT Devices
by: Fei, Hongming, et al.
Published: (2024)
by: Fei, Hongming, et al.
Published: (2024)
Monotonicity as an Architectural Bias for Robust Language Models
by: Cooper, Patrick, et al.
Published: (2026)
by: Cooper, Patrick, et al.
Published: (2026)
Review of Case-Based Reasoning for LLM Agents: Theoretical Foundations, Architectural Components, and Cognitive Integration
by: Hatalis, Kostas, et al.
Published: (2025)
by: Hatalis, Kostas, et al.
Published: (2025)
Enhancing Mathematical Problem Solving in LLMs through Execution-Driven Reasoning Augmentation
by: Basarkar, Aditya, et al.
Published: (2026)
by: Basarkar, Aditya, et al.
Published: (2026)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
by: Zhao, Lepeng, et al.
Published: (2026)
by: Zhao, Lepeng, et al.
Published: (2026)
Attacking Delay-based PUFs with Minimal Adversary Model
by: Fei, Hongming, et al.
Published: (2024)
by: Fei, Hongming, et al.
Published: (2024)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
by: Muñoz-Ortiz, Alberto, et al.
Published: (2023)
by: Muñoz-Ortiz, Alberto, et al.
Published: (2023)
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
by: Hill, Brennen, et al.
Published: (2025)
by: Hill, Brennen, et al.
Published: (2025)
Similar Items
-
TRIZ Agents: A Multi-Agent LLM Approach for TRIZ-Based Innovation
by: Szczepanik, Kamil, et al.
Published: (2025) -
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
by: Zhang, Peiyan, et al.
Published: (2025) -
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025) -
MarkLLM: An Open-Source Toolkit for LLM Watermarking
by: Pan, Leyi, et al.
Published: (2024) -
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)